Artificial intelligence

AI Made Cloud Waste Expensive Again. The Fix May Start with Provisioning

AI Made Cloud Waste Expensive Again

As AI pushes cloud spending higher, enterprises are finding that better dashboards may not be enough. The next battle over cloud efficiency may start much earlier, when infrastructure is defined and provisioned.

Idle GPU capacity is expensive, and increasingly difficult to ignore as AI drives demand for costly compute while cloud utilization remains stubbornly low.

An April 2026 analysis from cloud optimization company Cast AI, covering tens of thousands of production Kubernetes clusters across Amazon Web Services, Microsoft Azure and Google Cloud, reported average GPU utilization of just 5%. CPU utilization averaged 8%, down from 10% a year earlier, while CPU overprovisioning, the gap between capacity requested and capacity consumed, jumped from 40% to 69% year over year.

Flexera’s 2026 State of the Cloud Report, based on 753 cloud decision-makers, points to the same problem. Estimated waste across infrastructure and platform services rose to 29% of cloud spending, reversing five consecutive years of improvement.

AI is changing the economics of that waste. Excess CPU capacity could often be treated as something to optimize later. Idle accelerators spread across teams, workloads and regions make poor provisioning decisions considerably more expensive, shifting attention from finding waste after deployment to preventing it when infrastructure is first defined.

AI turns FinOps into a technology-wide discipline

The FinOps Foundation’s 2026 State of FinOps survey shows how quickly the landscape has changed. Among 1,192 practitioners responsible for more than $83 billion in annual cloud spending, 98% reported managing AI-related costs, up from 31% two years earlier. With 90% also managing SaaS costs, cloud cost management is increasingly becoming technology cost management.

But optimizing after deployment has limits. Extra capacity requested as a safeguard can become expensive when oversized configurations are embedded in reusable templates and replicated across applications, regions and teams. By the time a dashboard identifies the waste, the pattern may already be widespread.

That is shifting attention to an earlier control point: provisioning.

Platform engineering moves closer to the centre

The rise of platform engineering is happening at almost exactly the same time. Gartner projects that 80% of large software engineering organizations will have platform engineering teams by the end of 2026, up from 45% in 2022, and Google’s 2025 DORA research found 90% of organizations had adopted at least one internal platform.

That shift matters beyond developer productivity. Internal platforms increasingly shape how infrastructure is requested, configured, secured and deployed. Reusable components and standardized configurations give organizations a chance to address inefficient provisioning before it reaches the cloud bill. DORA also described AI as an amplifier of existing strengths and weaknesses, making platform quality increasingly important as AI adoption grows.

Enterprises are therefore standardizing infrastructure creation just as provisioning mistakes become more expensive. The question is whether that standardization produces measurable improvements in speed, consistency and cost, something a recent controlled multi-cloud experiment set out to examine.

A controlled experiment puts numbers behind automation

Evidence comes from a study presented at the 3rd World Congress on Smart Computing (WCSC2026) in Bangkok, Thailand. “Evaluating Terraform for Reproducible and Scalable Multi-Cloud Infrastructure Automation”, examined whether declarative Infrastructure as Code could improve provisioning speed and consistency across cloud providers.

The researchers built a controlled three-tier environment spanning networking, auto-scaled compute and managed databases, then deployed equivalent configurations on AWS and Azure across three replications. The experiment targeted a familiar platform-engineering problem: reproducing infrastructure reliably while multiple engineers continue to change it.

Vinoth Punniyamoorthy led the multi-cloud automation framework and modular Terraform design, while Srivenkateswara Reddy Sankiti contributed to performance benchmarking and statistical validation. Nachiappan Chockalingam worked on Git and Jenkins-based CI/CD integration, Bikesh Kumar developed the JMeter workload scenarios, and Aswathnarayan Muthukrishnan Kirubakaran contributed to the architecture and remote-state design using Amazon S3 and DynamoDB locking.

Nitin Saksena and Akash Kumar Agarwal of Albertsons Companies brought an enterprise architecture perspective to infrastructure sprawl, configuration drift and cost governance, while Seema G. Aarella of Austin College contributed to the statistical analysis and broader DevOps and cloud-governance framing.

On AWS, Terraform provisioned the environment in an average of 12.6 minutes, compared with 18.7 minutes for AWS CloudFormation and more than 42 minutes for manual provisioning. On Azure, Terraform averaged 13.4 minutes. The researchers attributed part of the difference to Terraform’s dependency graph, which allows independent components to be provisioned concurrently.

“Multi-cloud is the reality for most enterprises today. The question isn’t whether to automate, it’s whether your automation strategy is reproducible and scalable,” said Punniyamoorthy.

The cost finding is really about discipline

Over 30 days, the Terraform-managed environments cost $112.40 on AWS and $118.90 on Azure. The study reported the AWS result at roughly 13% to 15% below the CloudFormation comparison and close to 25% below manual provisioning, with an R² of 0.84 between the degree of automation and operating cost.

Those figures need context. Terraform does not make an identical cloud resource cheaper. The difference lies in the practices surrounding automation: reusable modules standardize configurations, automated teardown reduces forgotten resources, and version control makes inefficient definitions easier to identify before they spread.

The finding, then, is less about Terraform pricing than infrastructure discipline.

“In production environments, the cost implications of infrastructure sprawl and configuration drift are very real,” said Saksena. “Terraform’s module reuse and state management directly address those problems in ways that manual processes simply can’t replicate at scale.”

Configuration drift may be the more important finding

Another result may matter more to platform teams. The researchers simulated two engineers changing infrastructure simultaneously, a routine scenario in shared environments. With remote state in Amazon S3 and DynamoDB locking, drift incidents fell by close to 70%, while average conflict-resolution time dropped from 22.4 seconds to 6.8 seconds.

Under synthetic load of 1,000 requests per second, Terraform also produced a 25% to 30% faster auto-scaling response than the CloudFormation configuration tested. The researchers reported the differences as statistically significant at p < 0.05 using ANOVA across three replications.

Infrastructure code can be reviewed like application code, but the deployed environment remains shared. Without coordinated state, engineers can make changes based on different assumptions about what is actually running.

“Infrastructure as Code is no longer just a developer convenience, it’s a strategic governance tool,” said Aarella. “This research gives organizations empirical grounding to make informed decisions about how they automate and manage cloud infrastructure.”

Security increasingly overlaps with that problem. Google Cloud’s Threat Horizons reporting attributed 29.4% of initial-access incidents in the first half of 2025 to misconfiguration, falling to 21% in the second half as other attack vectors grew. Versioned definitions, controlled state and automated policy validation can therefore serve as governance controls as well as operational tools.

The ecosystem has changed too

Enterprises evaluating Terraform also face a question outside performance benchmarking: licensing.

HashiCorp moved Terraform from the Mozilla Public License to the Business Source License in 2023, making version 1.5.7 the last MPL release. The change prompted the creation of OpenTofu, a community fork later adopted by the Linux Foundation. IBM completed its $6.4 billion acquisition of HashiCorp in February 2025.

For enterprises standardizing infrastructure platforms over several years, licensing and vendor strategy now sit alongside performance and interoperability in the decision.

Cloud efficiency is becoming a design problem

Early cloud efficiency focused on visibility: finding idle resources and eliminating obvious waste. AI is changing that equation by making poor provisioning decisions considerably more expensive.

The study offers a narrower finding: standardized infrastructure definitions can improve reproducibility, while controlled state can help teams manage shared environments more consistently. Although the experiment did not test AI workloads, the same provisioning discipline could become increasingly relevant as enterprises deploy more costly GPU infrastructure.

As infrastructure becomes easier to provision, it also becomes easier to overprovision at scale. The next phase of cloud efficiency may therefore begin before deployment, making cost and resource discipline part of infrastructure design rather than something discovered later on a dashboard.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This