← Back to list

Tokenmaxxing and the failure of simplistic AI metrics

Earlier this year, a new term entered the tech lexicon: tokenmaxxing: the foreseeable result of businesses assessing their employees’…

Enrique Dans in Enrique Dans · 2026-05-31 10:55 · 127 claps · 4.9 min read paywalled
#ai #artificial-intelligence #tokenmaxxing #metrics #corporate-innovation
Open on Medium ↗
Wiki topics: AI · AI · General

Tokenmaxxing and the failure of simplistic AI metrics

Earlier this year, a new term entered the tech lexicon: tokenmaxxing: the foreseeable result of businesses assessing their employees’ performance on the basis of how many AI tokens they were using, how many calls, how many times agents acted as intermediaries and how much context they moved through their systems. Comfortable, quantifiable, easy to log on a dashboard and, above all, it provided the reassuring illusion of control that organizations like so much when a new technology appears.

This week’s Financial Times story about Amazon eliminating an internal ranking of AI use after some employees were maxxing token use to move up positions, is almost too perfect to be true. Kirorank scored users of the Kiro platform based on their AI activity, until some workers began assigning unnecessary tasks to autonomous agents. The company eventually had to tell employees something that should have been obvious from the beginning: don’t use AI just for the sake of it.

It’s hard to find a better example of Goodhart’s law: when a metric becomes a target, it’s no longer a good metric. But in this case we have to go a step further: token consumption did not cease to be a good metric when it became a goal, because in reality, it was never a good metric. It was, at best, a lazy proxy for activity. And activity is not value.

An employee who consumes a lot of tokens is not necessarily working better. They may be formulating their questions poorly, sending unnecessary context as if there were no tomorrow, using agents for trivial tasks, iterating without judgment, accepting mediocre answers, or delegating to the machine processes that he would have solved faster with a conversation, a search, or five minutes of concentration.

The AI industry has every incentive to drive tokenmaxxing. For companies that sell infrastructure, more inference means more demand. The narrative of agentic automation requires more loops, more tool calls, more memory, and more context seem like symptoms of sophistication. But for the company that pays the bill, the analysis should be exactly the opposite: more consumption does not mean more intelligence. Most of the time it means worse architecture.

Smart companies should not celebrate their systems consuming more and more, and instead ask why. Anthropic’s guide on context engineering formulates it quite clearly: the goal is to find the smallest possible set of high-signal tokens that maximizes the probability of obtaining the desired result. Not the largest set. Not the longest prompt. Not the most showy conversation, simply the smallest and most relevant set.

Anthropic’s guidelines should be framed in the boardroom of every company looking measure AI adoption among its employees. Because measuring tokens is easy. Measuring competition is much more difficult. A good professional can use few tokens because he knows exactly what to ask for, what context to provide, which model to choose, when to stop, and even when not to use artificial intelligence. A bad one, on the other hand, can use millions because they don’t know how to think about the problem, don’t know how to structure information, don’t know how to evaluate the answer or have learned that the dashboard rewards noise. In this scenario, the ranking does not identify the best users: it identifies the most expensive.

The paradox is uncomfortable: the truly competent employee may seem less “adoptive” than the one who turns every task into an unnecessary twenty-step agentic liturgy. The first one does engineering. The second does theater.

It is not a new phenomenon. Organizations have been destroying good intentions for decades through poorly chosen indicators: answered calls, lines of code, billable hours, number of meetings, closed tickets, publications, appointments, leads, visits, clicks. The same thing always happens. A metric is chosen first because it seems to correlate with something important. Then it becomes a target.

With AI, the problem is even more dangerous because the marginal cost of faking activity can be very high. An agent can run loops, call tools, retry, summarize, query documents, generate code, discard it, and start over. From the outside, everything looks like work.

That is why contrary signals are so important. METR’s study of experienced developers, for example, found that using AI tools made them take 19% longer to complete tasks on repositories they knew well, even though they themselves believed they were taking less time. This demonstrates that the subjective perception of productivity can be deeply misleading. And if perception is deceiving, a token counter is even more so.

That’s also why techniques such as OpenAI’s prompt caching, which can reduce latency and costs in repeated prompts, or Microsoft’s recommendations on RAG chunking, which insist on sending relevant information and eliminating the irrelevant, make sense. All these practices are based on the same idea: the token is not a medal, it is a resource. And like any resource, it must be administered.

The actual adoption of AI should not be measured by how much is consumed, but by how much the work improves. Less time to a correct decision. Fewer errors. Fewer unproductive reps. Better documentation. Better maintainable code. Better customer service. Better organizational learning. Better ability to address problems that could not be addressed before. And, above all, a better relationship between the result obtained and the cost incurred.

Clearly, the numerator matters, but so does the denominator: A company that only looks at tokens is measuring the denominator and pretending that it tells them something about the numerator. It’s like evaluating a driver by the liters of gasoline consumed, a researcher by the number of PDFs opened, or a teacher by the megabytes downloaded to prepare for a class. There may be some weak correlation in certain contexts, but it would be foolish to make it a performance criterion. The relevant question is not who uses more AI, it’s who gets the best results because they know when, how, and what to use it for.

This leads us to a fundamental distinction: access to inference capacity can become a very relevant part of the value proposition for certain professionals, as I proposed when talking about tokens as a form of remuneration or capacity for action. But it is one thing to equip a person well so that they can work better, and quite another to reward them for exhausting the budget. Giving access to powerful models can be an investment. Encouraging its indiscriminate consumption is accounting stupidity.

Business maturity in AI is not be about the millions of tokens processed: it’s about designing systems that need fewer tokens to achieve better results. Less brute force and more well-selected context. Fewer rankings and more criteria.

The Amazon episode should be an early warning. The company tried to accelerate a technological adoption through a visible, comparable and apparently objective metric. The problem is that people do not obey the abstract objectives of management: they obey the real incentives of the system. And if the system rewards tokens, they will produce tokens.

AI needs metrics, but very specific ones that capture value, quality, learning, reliability, safety, total cost, and real process improvement. It needs audits, benchmarks, controlled experiments, and discipline. It needs, in short, management.

Because when token consumption becomes a target, it stops measuring adoption. And when a company believes that token consumption measures intelligence, what it is really measuring is its own naivety.

(En español, aquí)


메타데이터
post_id
ee81820740b2
slug
tokenmaxxing-and-the-failure-of-simplistic-ai-metrics-ee81820740b2
url
https://medium.com/enrique-dans/tokenmaxxing-and-the-failure-of-simplistic-ai-metrics-ee81820740b2
canonical_url
https://medium.com/enrique-dans/tokenmaxxing-and-the-failure-of-simplistic-ai-metrics-ee81820740b2
author_url
https://medium.com/@edans
status
ok
fetched_at
2026-06-10 08:17:25