The Four meters: AI token unit economics now decide return on AI investments.
As the price of intelligence collapses with use of AI enabled solutions, enterprises must use AI token unit economics, not project budgets…
The Four meters: AI token unit economics now decide return on AI investments.
As the price of intelligence collapses with use of AI enabled solutions, enterprises must use AI token unit economics, not project budgets, to govern whether AI returns a profit, and I think four ‘meters’ will be key to decide the success here.
On Monday 27 January 2025, the most valuable company on earth lost about $589bn of market value in a single session, the largest one day fall in stock market history [Source: Bloomberg; CNBC, 27 January 2025]. Nvidia’s shares dropped roughly 17% and the trigger was an AI model!
Little known Chinese laboratory, DeepSeek, had released a system that rivalled the frontier on hard reasoning tasks while costing a fraction of the usual sum to train and serving answers far more cheaply than the incumbents [Source: Bloomberg; Yahoo Finance, January 2025]. Perhaps, this is a case of market repricing AI token unit economics!
27/1 looks like a story about chips and Chinese competition but I don’t think it is. The market repriced one number i.e. the cost of producing intelligence and assumed the rest would follow. The ‘cost of making a token’ and the ‘value a business’ draws from it are different quantities, and the second did not move that day at all. Nvidia’s own founder, Jensen Huang, called DeepSeek’s work “an excellent AI advancement” [Source: Nvidia statement, January 2025]. The statement essentially talks about supply, and not about value…. which was telling.
If the markets take the lesson of the day as ‘price of intelligence collapsing’, the task maybe is to ride it down and get the AI bill lower… which I think probably is the wrong lesson.
The right question is to ask on the right AI token unit economics of a business, and the four numbers that govern them needs managing?
Two findings from last year sit awkwardly side by side. The first is that intelligence has become radically cheaper. By the Stanford AI Index, the cost of querying a model at GPT-3.5 level on a standard benchmark fell from about $20/million tokens in November 2023 to roughly 0.07 dollars by October 2025, a decline of more than 280 times in under two years [Source: Stanford AI Index, 2025]. Andreessen Horowitz has described the pattern as inference cost falling about tenfold a year, quicker than the collapse of compute prices during the personal computer era [Source: Andreessen Horowitz, “Welcome to LLMflation”, 2025]. Epoch AI, measuring the price to reach a fixed level of capability, found declines ranging from 9 to 900 times a year depending on the task [Source: Epoch AI, “LLM inference prices have fallen rapidly but unequally”, 2025].
The second is that almost no one is capturing the windfall. MIT’s Project NANDA, analysing 300 deployments, found that 95% of generative AI pilots delivered no measurable profit and loss impact, against $30 to $40bn of enterprise investment [Source: MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025”, July 2025]. Gartner in 2025 has a forecast that a large share of generative AI projects will be abandoned after the proof of concept stage, citing poor data quality, weak controls, rising costs, and unclear value.
Put the two together and the symptom is plain i.e. spending is up, intelligence is cheap, returns are missing. The cause is less obvious, and The conventional explanation is that the models are not good enough yet, or the use cases are immature. Neither survives contact with the data, because the models behind the failed pilots are the same models behind the few deployments that do pay. What separates the two groups is not the technology but whether the organisation treats AI as a project with a budget or as a production system with a unit cost. A project is approved once, checked at a milestone, and forgotten whilst a unit cost is watched every day, because it moves every day. And here the cheapness works against i.e. the cheaper a token gets, the more tokens a capable system tends to consume, so the bill that was meant to fall quietly climbs instead.
AI token unit economics, the cost and the value of each token at the point of use, have replaced the project budget as the unit of account for enterprise AI, and the firms should model it as ‘cost per bit or a cost per query’ i.e. as a margin, and perhaps not a milestone.
Yes, this sounds like a finance point but is really an operating one. A margin is a ratio, and a ratio has a numerator and a denominator that both move. Manage only the denominator, the price you pay per token, and you get ‘a Klarna’ i.e. a low cost to serve and a falling quality of outcome. Manage only the numerator, the value created, and you get the 95% i.e. a convincing demonstration and a bill no one modelled for.
Having the discipline that produces a return…hold both in view at once, continuously, at the level of the individual task rests on four numbers.
The four meters of AI token unit economics
I have watched this play out across engagements over the past year, and clients that keep their economics intact tend to use/read the same four meters. Let’s borrow the image from a utility control room i.e. four dials on the wall, each one a ratio, each telling you something the others cannot. Rate, burn, yield and curve.
The word that matter is meter, not lever. A lever is something you pull once and walk away from whilst a meter is something you read continuously, because the thing it measures keeps moving. AI cost is a meter problem, and not a lever problem, which is why a one off optimisation project never holds. The four numbers below drift every time a model updates, a prompt changes, or adoption widens, so the discipline is not to optimise them once but to keep all four in view at the same time.
Rate is the price you pay per token, blended across models, caching and routing, not the sticker price on a vendor’s page. The shift is from sticker to blended i.e. from “we run on model X” to “we pay a known blended cost per thousand tokens after routing and cache”. The proof that this is a real meter, and not housekeeping, is in the routing research. RouteLLM, a 2024 study from researchers at UC Berkeley, Anyscale and Canva, showed trained routers cutting cost by about 85% while holding roughly 95% of GPT4 quality, simply by sending easy queries to cheaper models [Source: Ong et al., RouteLLM, 2024]. DeepSeek entered the market in January 2025 priced roughly 90% below comparable Western models, resetting everyone’s rate overnight (Source: Google). Caching belongs in the same meter i.e. on the major providers, a cached read of a stable prompt prefix costs about a tenth of a fresh input token on Anthropic and around half on OpenAI, so an agent that resends the same system prompt and tools on every step can cut a large slice of its input cost for no loss of quality [Source: Anthropic and OpenAI pricing documentation, 2026]. To measure this, can a team state the blended cost per thousand tokens for your largest use case, or only the list price of the model it runs on?
Burn is how many tokens a task consumes to get the job done. The shift is from “intelligence is free” to a metered budget per task. The arithmetic is unforgiving i.e. the cost of a single interaction is the input tokens multiplied by the input rate, plus the output tokens multiplied by the output rate, and output tokens typically cost four to five times more than input [Source: provider pricing documentation, 2026]. That asymmetry means a verbose model that rambles to an answer can cost several times a concise one for identical work. As the per token price falls, total consumption rises faster, and reasoning models sharpen the effect, because a short visible answer can sit on top of thousands of hidden internal reasoning tokens [Source: NVIDIA inference guidance 2025]. Klarna is its own evidence I.e. the assistant that did the work of 700 agents at launch was reportedly doing the work of around 853 by late 2025 as adoption widened [Source: Google]. To measure, set a token budget per task and alert when it breaks, the way one would a cloud spend alert.
Yield is the value created per token spent, expressed in the business outcome i.e. a resolved ticket, a settled claim, an approved loan, and not in tokens or API calls. The shift is from activity metrics to outcome metrics and the clearest proof is in how the best vendors now price. Intercom sells its Fin agent at $0.99 per resolution, not per conversation and not per token, and grew the product from $1mn to more than $100mn in annual recurring revenue on that basis, resolving over a million issues a week [Source: SaaStr, February 2026]. Set against a human interaction at roughly $6, a $0.99 resolution is an 83% reduction on every issue solved [Source: fin.ai]. To measure, can one state value per token for top use case?
Curve is how the bill behaves as adoption scales, whether unit cost falls with volume or the total runs away from you. The shift should be to move from a linear assumption to an engineered curve. Midjourney is the cleanest illustration i.e. it moved image generation off NVIDIA GPUs onto Google TPUs and cut monthly inference spend from about $2.1mn to under 700,000 [Source: Midjourney TPU v6e migration figures via Google]. Salesforce, unsure which curve its customers wanted, shipped three Agentforce pricing models in eighteen months with $2 a conversation at launch, then $0.10 an action, then a per user licence at $125 a month [Source: SaaStr, February 2026]. To measure, when usage rises tenfold, does one’s unit cost fall, or does your bill rise tenfold.

The token economy, in working terms
Before the four meters become second nature, it helps to see the machinery underneath them, because most of the confusion about AI cost may come from not knowing what is being bought and sold.
A token is a fragment of text, roughly four characters or three quarters of a word, that a model reads and writes in. A page of prose is about a thousand tokens and every model prices its work in tokens, charging separately for the tokens you send in (the prompt, the context, the instructions) and the tokens it generates back. The token matters because it is the smallest thing you pay for and, when the work lands, the smallest thing that creates value. It is the natural unit of account, in the way a kilowatt hour is for electricity.
On the supply side, the price of a fixed level of capability is falling fast, driven by better algorithms, cheaper hardware, and competition of the kind DeepSeek unleashed in January. On the demand side, enterprises are pushing far larger volumes through far more ambitious workloads. The result is a market where the unit price keeps dropping while the total bill keeps climbing, which is why a falling token price is not, on its own, good news.
The arithmetic is simple and worth understanding. The cost of one interaction is the input tokens multiplied by the input rate, plus the output tokens multiplied by the output rate. Output is the expensive half, typically four to five times the input rate [Source: provider pricing documentation, 2026], so a model that answers at length costs far more than one that answers tightly. Two discounts change the picture i.e. caching, which lets you reuse a stable block of context at roughly a tenth of the price, and batching, which trades immediacy for a lower rate.
A single question and answer is cheap and an agent is not. Agents loop, rereading their instructions and their growing history on every step, so token consumption can rise with the square of a session’s length. Reasoning models add a further layer, working through thousands of internal tokens before they produce a short visible answer. This is why agentic adoption, not chat, is what strains budgets, and why the pricing of agentic products is moving away from the token altogether.
The commercial layer is climbing a ladder and It began with per seat licences, the software industry’s default. It moved to per token and per conversation metering as usage decoupled from headcount. It is now reaching per outcome pricing, where you pay for a resolved result rather than the work that produced it, as with Intercom’s charge per resolution or Sierra’s outcome only model. Each rung ties price more tightly to value, and each exposes the seller to more risk if the product does not work.
The levers are well understood and they compound i.e. route easy queries to cheaper models, cache stable context, right size the model to the task, compress prompts, cap output length, batch what is not urgent, and choose infrastructure suited to the workload. Combined, published cases routinely report 60% savings with no measurable loss of quality [Source: Ong et al., RouteLLM, 2024; Chen et al., FrugalGPT, 2023].
The strategy should be to read the four meters together and to instrument them before you scale, not after the invoice arrives.
The pattern we may miss
The difference is generally not budget, sector or model but the sequence. The firms that hold their economics together, instrument the meter before they scale the use case.
This sounds like a small operational nicety but is the deepest point in the piece, because of what a missing meter does to behaviour. An organisation generally optimises the numbers it can see. Give it a clear cost figure and no clear value figure, and it will drive cost to the floor, because that is the only feedback it has, and it will do so even when value is quietly collapsing underneath. That is not a failure of intelligence or intent but is what any system does when only one of its gauges is lit.
So the structural observation is that the unit economics that cannot be seen, one will optimise badly, because the visible number becomes the only number. The winners are not smarter about AI but are better instrumented before they are big, and being well instrumented changes what the organisation is even able to optimise for.

For decisions….
The first implication concerns where the next pound goes. The instinct is to fund another pilot but the better move is to fund instrumentation i.e. put a live cost and value meter on each AI workload already in production before approving any expansion. In practice, this means per task token budgets with alerts when they break, spend attributed by team and use case rather than buried in one cloud invoice, and a dashboard that shows yield, not just consumption. None of it is exotic but could be the same observability discipline that cloud finance teams built a decade ago, being now applied to tokens. The question shifts from “what is our AI budget” to “what is our blended cost and our yield, per use case”.
The second concerns a habit to stop i.e. most enterprises route every request to their most capable model by default, which is the single most expensive thing they do. Right sizing and routing are not advanced engineering, and the published research puts the savings between 50 and 85% at near parity on suitable tasks [Source: Ong et al., RouteLLM, 2024; Chen et al., FrugalGPT, 2023].
The third concerns what we measure i.e. retire the activity dashboards, seats provisioned, calls made, adoption rates, and replace them with yield…which includes value per token, or value per outcome, in the currency of the business.
The fourth concerns the conversation with management. AI has been presented as a capital project with a return that arrives later so reframe it as a margin line that is live now and ask for the four meters by use case. A management that learns to ask for rate, burn, yield and curve will fund the second group of companies rather than the first.
In Summary,
Return to that Monday in January, and the $589bn dollars the market erased in an afternoon. I think the traders were right about one thing and wrong about everything that mattered. I surmise, they were right that the cost of producing intelligence had collapsed, and that DeepSeek had proved it in public…but they were wrong to treat that as the end of the story, because a falling price of production says nothing about whether anyone captures the value. The market had just repriced a single meter and mistook it for a reckoning.
For the enterprises watching, the lesson is the one the 95% keep learning the hard way. The price of intelligence falling is not the same as the value of intelligence rising, and the gap between those two is exactly where AI token unit economics live.
메타데이터
- post_id
- da29952ccd3e
- slug
- the-four-meters-ai-token-unit-economics-now-decide-return-on-ai-investments-da29952ccd3e
- url
- https://medium.com/@RaxRoshan/the-four-meters-ai-token-unit-economics-now-decide-return-on-ai-investments-da29952ccd3e
- canonical_url
- https://medium.com/@RaxRoshan/the-four-meters-ai-token-unit-economics-now-decide-return-on-ai-investments-da29952ccd3e
- author_url
- https://medium.com/@RaxRoshan
- status
- ok
- fetched_at
- 2026-07-13 06:23:13