Efficient to a Fault — Jevons Paradox in the AI Era
In 1865, a British economist named William Stanley Jevons published a book called The Coal Question and made an observation that nobody…
Efficient to a Fault — Jevons Paradox in the AI Era
In 1865, a British economist named William Stanley Jevons published a book called The Coal Question and made an observation that nobody wanted to hear. More efficient steam engines, he argued, would not reduce Britain’s coal consumption. They would, instead, increase it.
When the energy required per unit of work falls, the cost of that work falls with it, and falling costs don’t conserve resources; they build more factories, run more engines, and expand total demand faster than technical efficiency ever improves.
One hundred and sixty years later, Satya Nadella invoked this by name when DeepSeek briefly panicked Wall Street. He meant it as a growth thesis. He was also, without quite intending to, naming the exact mechanism that explains why AI’s efficiency story and AI’s energy story keep moving in opposite directions, and why both can be true at the same time.
The numbers that don’t add up
The technical data and the global energy reports appear to tell two different stories. DeepSeek reported spending approximately $5.576 million to train its V3 model, a figure that covers only the final training run and excludes prior experiments, hardware acquisition, and infrastructure costs that independent analysts at SemiAnalysis placed well above $500 million in total.
Yet, even the disputed version of that number pointed to something real: frontier-level performance achieved with fewer resources than Western labs had considered necessary. But while the efficiency story was still making headlines, the International Energy Agency was recording something else entirely. Data center energy consumption grew 17% in 2025, while global electricity demand grew 3%.
Capital investment from technology firms exceeded $400 billion that year and is projected to grow by a further 75% in 2026. By 2030, total data center consumption is expected to double, while AI-specific consumption is projected to triple. Per-query efficiency is improving. Total resource consumption is not.
What actually happens when something gets cheaper to use
This is not a contradiction; it is the rebound effect in operation.
When the cost of computation falls, organizations that previously couldn’t justify frontier-scale AI enter the market, existing users expand their workloads, and developers unlock applications that simply weren’t viable at higher price points. The Stanford 2026 AI Index illustrates how this plays out even within a single model comparison.
Despite its training cost efficiency claims, DeepSeek V3 consumes approximately 23 watts per medium-length prompt at inference, the phase that accounts for the dominant share of long-term operational costs. Claude 4 Opus consumes approximately 5 watts for the same task. The model celebrated for doing more with less uses 4.6 times as much energy when it actually does its job. And at the company level, the demand response to cheaper AI is unambiguous. As Dario Amodei wrote in the weeks after DeepSeek’s release: because the value of a more intelligent system is so high, efficiency gains cause companies to spend more, not less, on training models.
“The gains in cost efficiency end up entirely devoted to training smarter models, limited only by the company’s financial resources.”
There is a counterargument worth acknowledging. Some economists find that rebound effects in modern energy markets are modest enough that efficiency still delivers net consumption reductions, and the Jevons Paradox holds only when demand is highly price-elastic. But in a sector where Meta raised its annual AI spending to over $60 billion within weeks of a competitor publishing cheaper training figures, demand is not behaving like an inelastic market.

The measurement problem nobody is naming
This is where the efficiency story breaks down, in a way that goes beyond the paradox itself and into the measurement framework used to assess it. The metrics most commonly cited as evidence of AI sustainability are almost entirely supply-side rates. Power Usage Effectiveness (PUE) measures how efficiently a facility cools its servers relative to total power draw, but says nothing about absolute consumption. A data center can achieve a world-class PUE of 1.1 while its total electricity draw grows substantially year-over-year as server density increases.
Corporate sustainability reporting makes the same category error by focusing on training costs, a one-time event, while treating inference, which occurs billions of times daily, as a footnote. The Stanford 2026 AI Index documented what this produces: third-party estimates of training emissions for the same frontier model diverging by a factor of nearly two depending on methodology. The labs are not publishing audited figures. Academics are estimating from the outside.
The credible sustainability conversation needs to happen in the gap between those two numbers, and right now it largely isn’t. Sustainability is a function of volume, not rate, and the industry is measuring rates. To determine whether AI is becoming more sustainable, analysts need to model adoption curves, the price elasticity of AI services across sectors, second-order enabling effects in which AI unlocks consumption in adjacent industries, and lifecycle emissions that include hardware manufacturing and disposal.
That is demand modeling work, and it requires a different set of skills than infrastructure optimization. Both matter, but only one is currently at the table.
What Jevons actually tells us
Jevons was not arguing against efficiency. His point was that efficiency gains interact with demand in ways that intuition consistently gets wrong, and that treating efficiency as a conservation tool without accounting for the demand response will produce forecasts that miss the actual outcome every time. The AI industry is running that experiment live.
Every announcement of a more efficient model is followed, within months, by a capital expenditure figure that dwarfs whatever the efficiency gains were saved. The pattern is not incidental; it is the mechanism. None of this makes the engineering work pointless; it means that efficiency, reported in isolation, cannot tell us whether AI is becoming more sustainable. For that, we need the demand side of the equation: how much the market expands when costs fall, what new use cases get unlocked, and how total consumption moves relative to the gains.
Without that model, the industry will keep producing sustainability reports that look better each year while the grid gets more stressed. Jevons observed this while watching coal in 1865. The 2026 data is making the same argument. The question is whether the people writing the sustainability reports have read either of them.
메타데이터
- post_id
- cec1210e91cb
- slug
- efficient-to-a-fault-jevons-paradox-in-the-ai-era-cec1210e91cb
- url
- https://medium.com/@Xencee/efficient-to-a-fault-jevons-paradox-in-the-ai-era-cec1210e91cb
- canonical_url
- https://medium.com/@Xencee/efficient-to-a-fault-jevons-paradox-in-the-ai-era-cec1210e91cb
- author_url
- https://medium.com/@Xencee
- status
- ok
- fetched_at
- 2026-07-09 15:12:33