← Back to list

The Great AI Replacement Hit a Spreadsheet: Microsoft and Uber Can’t Afford Their Own Agents

In November 2025, Microsoft agreed to invest up to $5 billion in Anthropic, and Anthropic committed to buy $30 billion of Azure compute…

Kashif Mehmood in Towards AI · 2026-07-06 15:01 · 28 claps · 15.7 min read paywalled
#llm #artificial-intelligence #machine-learning #programming #technology
Open on Medium ↗
Wiki topics: LLM · Large Language Models AGT · AI Agents ML · Machine Learning AI · AI · General EDU · Education & Learning 💻 · Programming ☁️ · DevOps & Cloud

The Great AI Replacement Hit a Spreadsheet: Microsoft and Uber Can’t Afford Their Own Agents

AI agents were supposed to be the cheap option. The bill says otherwise.

AI agents were supposed to be the cheap option. The bill says otherwise.

In November 2025, Microsoft agreed to invest up to $5 billion in Anthropic, and Anthropic committed to buy $30 billion of Azure compute. Six months later, The Verge reported that Microsoft was removing most of its own engineers’ Claude Code licences, with the division that builds Windows and Microsoft 365 off the tool by 30 June. Microsoft still owns the stake. It still sells the models. It stopped paying retail for its own staff to use them, and that is the cleanest statement of the AI agents cost problem anyone has published: the companies that built the replacement can’t afford to run it on themselves.

I covered Zuckerberg telling Meta staff that agents are behind schedule; this is the companion problem, and the harsher one, because it holds even where the agents work. (internal link: Zuckerberg AI agents admission piece)

Key takeaways

  • Uber’s surging Claude Code use “maxed out” its full-year 2026 AI tools budget in roughly four months, per CTO Praveen Neppalli Naga (The Information, 14 April 2026). By June, Uber capped spend at $1,500 per employee per tool per month.
  • Gartner puts agentic tasks at 5 to 30 times the tokens of a chatbot interaction (March 2026). A Stanford-affiliated study measured the average agentic coding task at 4.17 million tokens and $1.857, against $0.023 for a coding chat (arXiv:2604.22750, April 2026).
  • Goldman Sachs forecasts token consumption rising 24x to 120 quadrillion tokens a month by 2030 as agents spread (May 2026).
  • Gartner separately warns spending on AI coding agents is on track to pass the average software developer’s salary by 2028.
  • Nvidia’s Bryan Catanzaro, April 2026: “For my team, the cost of compute is far beyond the costs of the employees.”

Microsoft kept the $5 billion stake and cancelled the licences

The details in Tom Warren’s Notepad report (The Verge, 14 May 2026) are worth reading closely, because each one cuts against the official framing. Microsoft opened Claude Code to employees in December 2025, and to more than engineers: project managers and designers got seats too. The tool spread fast, and by Warren’s account Microsoft’s own developers “favored Claude Code over GitHub Copilot CLI in recent months”. Then the Experiences + Devices division, the group behind Windows, Microsoft 365, Outlook, Teams and Surface, was told to wind down usage by the end of June.

Officially the move is standardisation. Rajesh Jha’s internal memo called Claude Code “an important part of that learning” while praising Copilot CLI as “a product we can help shape directly with GitHub”. Warren’s sources added the quieter half: “the decision is also a financial one. The June 30th cutoff is the last day of Microsoft’s current financial year, and canceling Claude Code licenses is an easy way to cut some operating expenses.”

Hold the two halves of the ledger side by side. The $30 billion Azure commitment from Anthropic, the up-to-$5 billion investment (Nvidia added up to $10 billion in the same announcement), the Foundry deal serving Claude to Azure customers: all untouched. The seats where Microsoft itself paid the per-token price of agentic coding: cancelled at fiscal-year end, exactly when a CFO goes hunting. A company that believed agents were about to be cheaper than engineers would have done the opposite trade.

Uber’s 2026 budget lasted four months

Uber ran the same experiment with a leaderboard on top, and produced the year’s most quotable budget disaster. On 14 April 2026, Uber CTO Praveen Neppalli Naga told The Information that surging use of AI coding tools, particularly Claude Code, had maxed out the company’s full-year 2026 budget for them. His words: “I’m back to the drawing board because the budget I thought I would need is blown away already.”

The curve behind that sentence, per the same report and Forbes’ follow-up (17 May 2026): Claude Code rolled out to roughly 5,000 engineers in December 2025. In February, 32% of engineers were using it. By March, 84% were classified as agentic coding users. A typical engineer cost $150 to $250 a month in tokens; power users ran $500 to $2,000. Naga reported burning about $1,200 in a single two-hour demo session. And Uber had built internal leaderboards ranking teams by AI tool usage, which converted a metered bill into a competition. Nobody wins a leaderboard by prompting less.

The reckoning arrived in stages. Uber’s president and COO Andrew Macdonald called the moment the CTO’s disclosure landed “head-exploding” (Rapid Response podcast, 23 May 2026), and told Fortune the productivity link was missing: “it’s very hard to draw a line between one of those stats and ‘Okay now we’re actually producing like 25% more useful consumer features’”. On 2 June, Bloomberg reported Uber capping AI tool spend at $1,500 per employee per tool per month.

Here is what makes Uber the perfect specimen rather than a cautionary outlier: the tools worked. Uber’s own Q1 2026 prepared remarks (6 May) state that “more than 10% of production-ready code” is “now driven autonomously by AI coding agents” and 95% of engineers use AI tools monthly. Real output, real velocity. The budget died anyway, and it died on arithmetic while the capability story stayed intact.

Tokenmaxxing turned the meter into a scoreboard

The Microsoft and Uber stories read differently once you see the culture that produced them. Through late 2025 and early 2026, the biggest employers in tech tied status, and sometimes performance reviews, to AI usage volume. Employees responded the way employees always respond to a metric.

At Amazon, the Financial Times reported (12 May 2026) that staff were using an internal agent tool called MeshClaw “to automate non-essential tasks in a bid to show managers they are using the technology more frequently”, after Amazon set targets for more than 80% of developers to use AI weekly and began tracking token consumption on internal leaderboards. “There is just so much pressure to use these tools,” one employee told the FT. The workers’ own slang for the behaviour: tokenmaxxing. Amazon said the statistics would not feed performance evaluations; several staff told the FT they believed managers were watching them anyway.

At Meta, an employee built a dashboard called Claudeonomics that ranked colleagues by token consumption across more than 85,000 employees, complete with “Token Legend” and “Cache Wizard” titles for the top 250 (The Information, 7 April 2026; details via Fortune, 9 April). Thirty-day usage on the board exceeded 60 trillion tokens. The top user averaged 281 billion tokens, which Fortune priced at potentially over $1.4 million at Claude Opus 4.6 rates. For one employee. The dashboard came down two days after the story broke (Meta says the employee took it down unprompted), and by mid-June The Information reported a Meta memo warning internal AI use was “on track to cost billions in 2026”, with a new AI Gateway dashboard adding budgets and spending alerts.

How Microsoft, Uber, Amazon, and Meta reined in exploding AI coding agent costs in 2026

How Microsoft, Uber, Amazon, and Meta reined in exploding AI coding agent costs in 2026

By 26 June, CNBC was describing the reversal as an era ending: “That led to the era of so-called tokenmaxxing and AI leaderboards, where employers have incentivised developers to use as much AI as possible without worrying about the results. The crackdown is underway.” D.A. Davidson’s Gil Luria told CNBC that some of the labs’ largest enterprise customers “may start limiting their out-of-control token spend”. Eighteen months of “adopt or be left behind”, and the new corporate virtue is switching it off.

Why AI agents cost more per task the harder they try

None of this is a procurement blunder, and that is the part worth understanding properly. Agent costs curve upward for mechanical reasons, and the mechanics are now measured.

Start with the multiplier. Gartner’s March 2026 inference-cost analysis states it flatly: “Agentic models, for example, require between 5–30 times more tokens per task than a standard GenAI chatbot.” For coding specifically, the multiplier is far worse. A study from Michigan, Stanford, MIT and Microsoft researchers (arXiv:2604.22750, April 2026) ran eight frontier models through SWE-bench Verified tasks with the OpenHands agent and measured what the meter actually reads. The average agentic coding task consumed 4.17 million tokens and cost $1.857. A multi-round chat about a coding problem: 3,390 tokens, $0.023. Roughly 1,200 times the tokens, for work in the same domain.

Agentic coding versus chat, measured. Source: Bai et al., arXiv:2604.22750, Figure 1.

Agentic coding versus chat, measured. Source: Bai et al., arXiv:2604.22750, Figure 1.

The shape of that spend explains why it resists optimisation. The study’s most useful number is the input/output ratio: 1.33 for chat, 153.85 for agents. Almost the whole bill is the context the agent re-reads; the code it writes is a rounding error on top. Every step of the loop (read the repo, plan, edit, run tests, read the failure, plan again) ships the accumulated history back through the model, so the fortieth step carries the freight of the thirty-nine before it. Input tokens dominate the cost “even when token caching is enabled”, the authors note. Late-loop tokens cost more than early ones by construction, which is why an agent that perseveres is an agent that gets more expensive per minute the longer it works.

Then add the failure economics. The best frontier models still fail roughly a third of SWE-bench Pro tasks, Scale AI’s contamination-resistant benchmark of hard, multi-file, real-repository fixes: Claude Opus 4.8 leads at 69.2%, Claude Sonnet 5 posts 63.2% (Anthropic, 30 June 2026). (The older SWE-bench Verified board is saturated above 70% and has been for a year; anyone still quoting “agents solve less than half” is quoting 2024.) A third of hard tasks failing would be fine if failures were cheap. They are the opposite. A failed attempt bills every one of its millions of tokens at full price, and the study found runs on the same task varying by up to 30x in total tokens, with accuracy that “often peaks at intermediate cost” and then degrades. Past a point, the extra spend buys unproductive exploration.

A plumber like this would not stay in business. He charges full call-out rate for every visit, including the four where the pipe still leaked, and his hourly rate rises the longer he stays because he re-inspects the whole house before touching the next joint. That is the agentic billing model, priced per token.

Why agentic coding gets more expensive the longer it runs: every loop re-reads the accumulated context

Why agentic coding gets more expensive the longer it runs: every loop re-reads the accumulated context

One more measured fact closes the trap: the models cannot tell you the bill in advance. The same study asked each model to predict its own token usage before execution and got weak-to-moderate correlations at best, with every model systematically underestimating. “Agents are not capable of predicting their own token costs,” co-author Jiaxin Pei put it. Humans quote a price before doing a job. Agents invoice you afterwards, and they guess low.

The machine stopped being the cheap part

The industry’s own principals have started saying the quiet arithmetic out loud. Bryan Catanzaro, Nvidia’s VP of applied deep learning research, told Axios in April 2026: “For my team, the cost of compute is far beyond the costs of the employees.” His employer’s CEO frames the same ratio as an ambition. On the All-In podcast at GTC in March, Jensen Huang said: “If that $500,000 engineer did not consume at least $250,000 worth of tokens, I am going to be deeply alarmed.” Half a salary in tokens, per engineer, as the target.

Hardware is attacking the cost side. Microsoft’s Maia 200 inference accelerator (announced January 2026: TSMC 3nm, 216GB of HBM3e) is the kind of silicon that bends the curve, and on the April earnings call Satya Nadella said it delivers “over 30% improved tokens per dollar compared to the latest silicon in our fleet”, live in Iowa and Arizona.

Maia 200, Microsoft’s inference chip. Source: Microsoft, official blog, 26 January 2026.

Maia 200, Microsoft’s inference chip. Source: Microsoft, official blog, 26 January 2026.

The catch sits in Gartner’s same March analysis. Inference on a trillion-parameter model should cost providers over 90% less in 2030 than in 2025, and Gartner still expects enterprise AI budgets to feel none of it, because agentic demand grows faster than unit cost falls and providers keep a slice. Analyst Will Sommer’s warning to product chiefs: they “should not confuse the deflation of commodity tokens with the democratization of frontier reasoning”. Or, as he told CIO Dive, “the customer isn’t going to see all of this money.” Gartner’s companion warning, via DevOps.com (30 June 2026), lands the punchline for engineering leaders: spending on AI coding agents is on track to pass the average software developer’s salary by 2028, and roughly 25% of tech leaders already spend $200 to $500 a month per developer on tokens.

The per-seat arithmetic, assembled from the verified figures in this piece, tells the story in one column:

The per-seat arithmetic of AI coding agent costs, 2026 to 2028

The per-seat arithmetic of AI coding agent costs, 2026 to 2028

Three years, three orders of magnitude, and the endpoint the industry’s biggest chipmaker publicly hopes for is half a salary per head.

Goldman Sachs modelled where the demand curve goes. Jim Schneider’s May 2026 report, “Decoding the Agentic Economy”, forecasts token consumption multiplying 24x between 2026 and 2030, to 120 quadrillion tokens a month, with enterprise agents lifting consumption 55x by 2040. Read the framing carefully: this is a bullish report. Goldman counts the explosion as a “margin inflection” for AI providers, because chips get 60–70% cheaper per token per year while usage grows faster still. Every quadrillion of that margin is somebody’s operating expense, and the somebody is the enterprises deploying agents.

Goldman’s Exhibit 2: consumption up more than 24x by 2030. Source: Goldman Sachs Global Investment Research, “Decoding the Agentic Economy”, May 2026, via ZeroHedge.

Goldman’s Exhibit 2: consumption up more than 24x by 2030. Source: Goldman Sachs Global Investment Research, “Decoding the Agentic Economy”, May 2026, via ZeroHedge.

None of this should surprise anyone who read the Wall Street Journal in October 2023. Tom Dotan and Deepa Seetharaman reported that GitHub Copilot, priced at $10 a month, was losing more than $20 per user per month on average in early 2023, with heavy users costing up to $80, according to a person familiar with the figures. GitHub disputed it (CEO Thomas Dohmke told Semafor that in November the per-user cost of goods sold was below the price). Either way, the industry absorbed the wrong lesson from the flat-rate era: that AI coding assistance costs roughly a lunch per month. In June 2026, GitHub itself moved Copilot to usage-based “AI Credits”, and Tech Times reported agentic users seeing roughly 10x cost surges. The subsidy is over. The meter is the product now.

Jim Covello, Goldman’s head of global equity research, asked the durable question back in June 2024 in “Gen AI: Too Much Spend, Too Little Benefit?”: “We estimate that the AI infrastructure buildout will cost over $1tn in the next several years alone… So, the crucial question is: What $1tn problem will AI solve?” His sharpest line then reads even sharper against the 2026 evidence: “Replacing low-wage jobs with tremendously costly technology is basically the polar opposite of the prior technology transitions I’ve witnessed in my thirty years of closely following the tech industry.” In his April 2026 update he marked his views to market and doubled down, per Fortune: “in this cycle, the chip companies are thriving at the expense of everyone above them”, and “Something has to change with this dynamic, either the companies higher in the chain need to start earning a return on investment or they will eventually need to spend less.”

Who is actually earning the buildout: net income growth since ChatGPT-3.5. Source: Goldman Sachs Global Investment Research, via Fortune, May 2026.

Who is actually earning the buildout: net income growth since ChatGPT-3.5. Source: Goldman Sachs Global Investment Research, via Fortune, May 2026.

The honest counter-case: a repricing is underway, and it has a floor

An argument this convenient deserves its strongest opposition, so here it is.

Token prices are collapsing, and fast. Epoch AI’s analysis of inference pricing found costs falling 9x to 900x per year depending on the capability milestone, with GPT-4-level performance dropping about 40x per year. a16z measured GPT-3-level capability falling 1,000x in three years, from $60 per million tokens in late 2021 to $0.06. Anthropic launched Claude Sonnet 5 in June at $2 per million input tokens. Betting against this curve has embarrassed everyone who tried.

The Copilot loss story preceded a real business. The product WSJ said was losing $20 a user became, by Nadella’s July 2024 earnings call, part of a GitHub run-rate of $2 billion, with Copilot driving over 40% of GitHub’s annual growth and crossing 20 million all-time users a year later. Early unit-economics horror stories sometimes describe a pricing experiment, and the experiment can succeed.

The output is real. Uber’s 10%-of-production-code figure sits in the company’s own earnings documents, alongside “thousands of updates deployed each week”. Anthropic said in May 2026 that more than 80% of the code merged into its production systems is written by Claude. The agents work. The dispute is about price, and prices move.

And the substitution valve is already open. When the closed-frontier bill explodes, teams do not go back to typing. They route the loop to open-weight models at a fraction of the rate card. Z.ai’s GLM-5.2 ships under MIT licence at $1.40/$4.40 per million tokens and beats GPT-5.5 on FrontierSWE, an independent long-horizon coding benchmark (74.4 vs 72.6), at roughly a sixth of the cost. (internal link: GLM-5.2 piece) Kimi K2.7 Code lists at $0.95/$4.00. MiniMax M3 runs $0.30–0.60 in and $1.20–2.40 out. (internal link: MiniMax M3 piece) The arbitrage is already live: CNBC’s June crackdown piece reports the startup Lindy moving 100% of its traffic from Claude to DeepSeek and saving millions. If the “wall” is a price, open weights are a door in it.

So the steelman says: 2026 is a repricing, the kind every young technology goes through, and the correct headline is “enterprise AI spend gets rationalised”, which happens to every line item eventually.

Take the steelman seriously and notice what survives it. Cheaper tokens per unit have coexisted with exploding bills at every step so far, because the multiplier (5–30x per task, ~1,200x for coding agents, 24x aggregate by 2030) has outrun the discount, and because the thing enterprises actually buy is frontier-model agentic reasoning, precisely the product Gartner says stays scarce and expensive. The Stanford study’s grimmest finding was that more tokens frequently buy worse answers past the intermediate-cost peak, which means the spend curve cannot be fixed by spending. And the substitution valve, real as it is, concedes the article’s thesis: companies are re-routing around the flagship agents because the flagship agents cost more than the humans were supposed to.

The role that survives is the one holding the off switch

Look at what every company in this story actually did after the shock, because they all did the same thing. Microsoft standardised on the tool whose costs it controls. Uber installed a cap. Meta built a budget-alert gateway. Amazon told staff the token stats would not feed reviews. Duolingo, which declared itself “AI-first” in April 2025 and made AI use part of performance evaluations, scrapped that metric by April 2026. CEO Luis von Ahn: “At the end, we backtracked, and we said, ‘No. Look, the most important thing in your performance is that you are doing whatever your job is as well as possible.’”

None of them turned the agents off. All of them put a human between the agent and the invoice. The logic follows straight from the mechanism above. An agent cannot predict its own costs, retries at full price, and gets more expensive the longer it persists, so the only reliable spend-limiter in the loop is the person who reads the output, judges whether attempt three deserves an attempt four, and kills the run that has started exploring. Call the role verification, review, oversight, whatever the org chart prefers. In 2026 it is the highest-leverage seat in the building, because every hour of it prevents a four-figure loop.

Even the maximalist vision now ships with a human at the centre. Huang’s line at GTC was “Every engineer is going to have a hundred agents”, and his 10-year sketch for Nvidia was 75,000 employees working with 7.5 million agents. A hundred to one, arranged around each person. Multiply Uber’s per-seat token bills by a hundred agents and you see exactly why the employee stays: someone has to be accountable for whether those hundred meters are buying anything.

Where AI agent costs overtook human costs on the capability curve, 2026

Where AI agent costs overtook human costs on the capability curve, 2026

For two years, the story was that your salary was the expensive line item and the digital worker was the cheap one. The 2026 numbers, from the companies most committed to the story, read the other way. A working engineer costs the same this month as last month, gets faster with experience at flat cost, quotes before working, and never re-reads the entire codebase at your expense before every keystroke. The machine that was meant to undercut that deal currently bills $1.86 a task, fails a third of the hard ones, and cannot estimate its own invoice. Humans remain the cheapest reliable thinking engine on the market: one runs on a sandwich and a night’s sleep.

The Great Replacement was not stopped by a protest. It was stopped by a spreadsheet.

Happy Coding ❤

Sources


메타데이터
post_id
958bfeeeeacd
slug
the-great-ai-replacement-hit-a-spreadsheet-microsoft-and-uber-cant-afford-their-own-agents-958bfeeeeacd
url
https://pub.towardsai.net/the-great-ai-replacement-hit-a-spreadsheet-microsoft-and-uber-cant-afford-their-own-agents-958bfeeeeacd
canonical_url
https://pub.towardsai.net/the-great-ai-replacement-hit-a-spreadsheet-microsoft-and-uber-cant-afford-their-own-agents-958bfeeeeacd
author_url
https://medium.com/@kashif-mehmood-km
status
ok
fetched_at
2026-07-08 20:12:56