Your GPU Bill Isn’t a Capacity Problem. It’s a Utilization Problem.
Enterprises are buying more GPUs to fix a shortage that a five-percent utilization rate says doesn’t exist.
Your GPU Bill Isn’t a Capacity Problem. It’s a Utilization Problem.

Enterprises are buying more GPUs to fix a shortage that a five-percent utilization rate says doesn’t exist.
Cast AI analyzed tens of thousands of Kubernetes clusters running across AWS, GCP, and Azure for its 2026 State of Kubernetes Optimization Report, and the number that should stop every CTPO mid-scroll is this: average GPU utilization came in at 5%. Not 50%. Five. Ninety-five percent of the GPU capacity enterprises are paying for, hour after hour, is doing nothing.
That finding lands in the middle of a year when, according to Gartner’s January 2026 forecast, spending on AI infrastructure (servers, network fabric, processing silicon) will add $401 billion in new spending, part of a worldwide AI spending total Gartner puts at $2.59 trillion, up 47% year over year. Boards are approving that spend on the premise that compute is scarce and every GPU secured is a GPU that keeps the roadmap alive. The Cast AI data says the premise is wrong for a meaningful share of that money. The GPUs aren’t scarce inside the walls where they’ve already been bought. They’re idle.
I’ve spent seventeen years watching organizations solve the wrong version of a capacity problem: first in SaaS scaling, then through a cloud migration where “we need more servers” turned out to mean “we’ve never turned off anything we provisioned for a traffic spike eighteen months ago.” GPUs are the same story with three more zeros on the invoice. And this time, because inference is metered per token and per second in a way legacy compute never was, the cost of not noticing shows up on next month’s bill instead of getting buried in a three-year depreciation schedule.
The scarcity story is easier to tell than the utilization story
“We can’t get enough GPUs” is a story that flatters everyone who tells it. It flatters the infrastructure team, because the problem is external — Nvidia’s allocation, the cloud provider’s capacity, the market. It flatters the finance team’s mental model, because more spend against a documented shortage is a defensible line item. And it flatters leadership, because chasing scarce compute sounds like urgency and ambition rather than an admission that nobody is watching what’s already running.
The scarcity story isn’t fabricated. Lead times on frontier accelerators are real, and pricing has moved against buyers: Cast AI’s own report notes that cloud vendors raised H200 pricing 15% this year, breaking a stretch of roughly two decades in which compute costs mostly fell. But scarcity at the point of acquisition and waste at the point of use are two different problems, and enterprises have spent 2025 and 2026 solving the first one by throwing more capital at it while leaving the second one almost entirely unmeasured. Cast AI’s president and co-founder, Laurent Gil, put the asymmetry plainly: “A GPU sitting idle costs dollars per hour. A CPU sitting idle costs cents.” When idle compute costs cents, nobody builds a dashboard for it. When it costs dollars per hour, not building that dashboard is a board-level oversight.
Where the five percent actually goes
The pattern is structural, not anecdotal, and it will look familiar to anyone who lived through the early years of cloud elasticity before autoscaling and spot markets matured. Teams provision GPU clusters for worst-case load: the demo, the launch spike, the quarter when the roadmap says usage triples, and then leave that capacity running at that size indefinitely, because rightsizing gets treated as a deployment-time decision instead of a continuous one. Cast AI’s report calls this out directly: configuration that was accurate six months ago is not accurate today, and treating a rightsizing pass as a one-time task is why the gap between provisioned and used capacity keeps widening instead of closing.
Training runs make the pattern worse, not better. A cluster gets hammered for two or three weeks during a training run, then sits mostly idle until the next cycle, because there’s no default mechanism, and often no incentive, to let anything else use that capacity in between. Multiply that across a portfolio of models, teams, and environments that each provisioned their own GPU pool “to be safe,” and 5% utilization stops looking like an anomaly and starts looking like the expected output of how enterprises buy and schedule compute today.
It isn’t only a GPU story, either, which is the detail that should worry finance the most. The same Cast AI analysis found CPU utilization at 8% and memory utilization at 20% across the same clusters in 2025. GPUs are simply the most expensive expression of a habit the organization already had. If the fix were “buy fewer GPUs,” the CPU and memory numbers wouldn’t tell the same story. The fix is rightsizing and scheduling discipline that Kubernetes was supposed to make automatic and, per the same data, largely hasn’t.
The counter-argument worth taking seriously
There’s a real case for provisioning ahead of need, and it deserves more than a token acknowledgment before getting waved off. If GPU procurement truly takes months and a competitor’s launch or a customer’s usage curve can move faster than your next allocation window, holding excess capacity is a hedge against a supply chain you don’t control, not a governance failure. Under-provisioning during a demand spike can cost more in lost revenue or reputational damage than the idle-GPU bill ever will, and a CTPO who strips buffers down to the theoretical minimum and gets caught flat-footed by a launch will not get credit for the clean utilization dashboard.
That argument holds for a slice of the capacity — a genuine strategic reserve, sized deliberately and revisited on a schedule. It stops holding once idle capacity is 95% rather than 20% or 30%, and once the same pattern shows up in CPU and memory pools that have never faced the acquisition constraints GPUs do. A hedge against scarcity is supposed to look different from business-as-usual overprovisioning; at 5% utilization, they’ve become indistinguishable, which is itself the evidence that this isn’t a considered hedge. Nobody sized a 95%-idle GPU fleet on purpose. They sized it once, walked away, and let metered billing keep charging for the decision every hour since.
The market is already telling you this is a management problem, not a supply problem
The clearest signal that this has shifted from an infrastructure question to a governance question is what’s happening inside FinOps teams. The FinOps Foundation’s 2026 State of FinOps report (a survey of 1,192 practitioners representing more than $83 billion in annual cloud spend) found that 98% of FinOps teams now manage AI spend, up from 63% in 2025 and just 31% in 2024. When a discipline built to instrument and govern cloud cost triples its AI coverage in two years, that’s an organization-wide acknowledgment that AI compute needs the same continuous scrutiny cloud spend eventually got, not a one-time procurement decision followed by silence. Tellingly, when the same survey asked FinOps practitioners what tooling capability they most wanted and don’t yet have, the top answer was granular monitoring of AI spend: tokens, LLM requests, and GPU utilization specifically. The people closest to the invoices are asking for visibility that, per Cast AI’s data, still mostly doesn’t exist.
The question that actually matters
“How do we get more GPUs” is the wrong question for most organizations asking it in 2026, because the honest answer for the majority of them is that they already have enough — they just can’t see what the ones they have are doing. The right question is narrower and less comfortable: what is our utilization curve right now, who owns it, and does anything in our stack rightsize continuously instead of once at deployment.
That question doesn’t have the urgency of a supply shortage. It has the tedium of an operating discipline, the same discipline cloud infrastructure eventually forced on organizations a decade ago, minus the grace period this time, because the meter is running per token instead of per instance-month. A CTPO who can answer it with a number and a name is managing an architecture problem. One who can only answer it with a procurement request is still managing a story.
Sameer is a CTPO with 17 years across SaaS scaling, cloud migration, and enterprise product leadership, including 12 years in OTT/media. He writes on AI infrastructure economics and engineering leadership at medium.com/@sameerkoli.
메타데이터
- post_id
- 324ef108e40d
- slug
- your-gpu-bill-isnt-a-capacity-problem-it-s-a-utilization-problem-324ef108e40d
- url
- https://medium.com/@sameerkoli/your-gpu-bill-isnt-a-capacity-problem-it-s-a-utilization-problem-324ef108e40d
- canonical_url
- https://medium.com/@sameerkoli/your-gpu-bill-isnt-a-capacity-problem-it-s-a-utilization-problem-324ef108e40d
- author_url
- https://medium.com/@sameerkoli
- status
- ok
- fetched_at
- 2026-07-19 00:45:20