← Back to list

Coffee, Tokens, and The Tokenpocalypse

My name is Yossi, also go by Josef, and I have a confession: I’m addicted to performance enhancing drugs.

Josef (Yossi) Goldstein in Wix Engineering · 2026-06-27 18:47 · 53 claps · 9.5 min read
#ai #token-economy #software-engineering
Open on Medium ↗
Wiki topics: AI · AI · General 🍳 · Food & Cooking

Coffee, Tokens, and The Tokenpocalypse

The Trojan Women Set Fire to their Fleet by Claude Lorrain (Italy 1643)

The Trojan Women Set Fire to their Fleet by Claude Lorrain (Italy 1643)

My name is Yossi, also go by Josef, and I have a confession: I’m addicted to performance enhancing drugs.

Namely: Caffeine and Tokens.

Lucky for me, my employer is gracefully willing to fund both of those habits.

For now.

When A Habit Becomes An Addiction

I remember how it started for me with Caffeine. I was in my late teens. We had in our high school one of those vending machines that spewed a horrible over-sweetened concoction that I used to pass for coffee. It was always available, it was $1 a cup, it gave me the boost and I was happy.

Over the years this became a morning routine. One cup became two, and then three. Before I knew it, without my cup o’ joe I couldn’t function as a human with a fully developed frontal lobe. Also at that point, the crappy coffee from the vending machine didn’t cut it anymore. I graduated to instant coffee, and then pods, and then high end espresso. Now suddenly my $1 fix that I got so used to, turned into $3 and then $5, several times a day. Just so I can be functional and productive at my job.

I did the math. Depending on how you price it, I consume over $300 in coffee cups a month. To my employer it costs much less of course. They buy the beans in bulk, they own the machine, they are employing the barista. But even if it’s only half that much, it still feels like a lot.

Yet we now have people around us that might spend the same sum in tokens in a week, and soon maybe in a day

Can’t start the morning without it..

Can’t start the morning without it..

We Burned The Ships

A few years ago, an LLM was a party trick. A fun thing you played around with, showed at a demo and everyone clapped. Then it became a productivity tool. Then somehow we all collectively decided (or it was decided for us) that it’s the tool we should use for everything.

So somewhere around 2025, the average developer stopped being someone who writes code and started being someone who directs an agent. The keyboard became less of an instrument of creation, and more of a proxy for communicating with a creative partner. A partner you now depend on and runs on tokens someone needs to pay for.

At first it was cute. “I asked Claude to write my unit tests.”. Then it became wholesale the way we write most of our code. Then it became the infrastructure to how we think through things and workflows to implement those ideas. Now as we are all getting more “AI Native”, there are entire companies whose engineering departments are, functionally, becoming more and more dependent on chains of skills and now agentic workflows. Churning tokens at a compounding rate.

And now? Engineering functions gradually can’t perform without it. Junior developers who joined the industry in the last two years have never really written a non-trivial piece of code without some model holding their hand. Product managers are writing PRDs by narrating loosely into a chat window. Designers are prompting their way through what used to be craft.

The cognitive musculature atrophies, while everyone is shipping features faster than ever.

Always one prompt away from delivering that ticket

Always one prompt away from delivering that ticket

We are all hooked, and we are definitely not going back. As with the proverbial conqueror who landed in the new world, we are burning the ships behind us to make this new world ours.

And all of this was fine and good, as long as our dependency remained economical…

The Numbers Don’t Add Up Anymore

The thing about addictions: they only feel sustainable until they don’t.

The article that you are reading right now is roughly 3000 tokens long. That means that having Claude read it would cost somewhere between $0.003 and $0.015, depending on the model, which is practically nothing.

Having one of those models generate this article, assuming it spent some tokens also thinking before writing, would cost somewhere between $0.03 and $0.15. So 10 times more, but still practically nothing.

What if Claude and I decided to work together on this article through about a dozen turns of conversation? Now we are talking more like $0.6 to $3, that’s not 10 times more, that’s x20 at this point. Add a harness prompt, a map of MCPs and skills, some tool calls… We are easily at $4-$14, maybe even more if Claude decided to get frisky. That’s almost my entire daily coffee budget, for one freaking article!

Now scale that up. From one article, to an engineering org.

Teams upon teams running agentic workflows through a full workday. They are not making a dozen turns per task. They’re running hundreds of them. Code generation, review loops, documentation generation and digestion, architecture brainstorms, ticket refinements. Each one chewing through context windows the size of small novels. Each one triggering tool calls that spawn sub agents and then more calls. And that’s before we even talked about the latest craze of loop engineering and dynamic workflows that offer a way to do all of that in a loop, or recursively, or both at the same time.

And for a while tech organizations were cool with it. There were bold claims by people like Jensen Huang, claiming engineers should be consuming token worth at the very least half their salary. Hell, for a short while some companies even measured people based on how many tokens they spent. That until the bills started arriving and “Tokenmaxxing” stopped being cool overnight.

Because all the data that we see right now, points out that the productivity simply doesn’t scale linearly with the money spent on tokens. Even more fundamentally, there’s a problem with Jensen’s thesis: If your $500k engineer is spending $250k annually on tokens, they better bring additional value to the company that translates to at least the same amount in revenue. And I don’t know about you, but I’m not seeing it, and neither are the quarterly earnings of those companies. Simply put, all signs are showing that the ROI simply isn’t there or at least has yet to be proven.

So if the added revenue is not there to fund those tokens, there are only two options left: You either curb the token spending or you fire someone so you can continue to fund the habit. Both feeling like an impossibility at this stage, but both are actually happening.

We are doing this because we care.. and because you are killing our budget.

We are doing this because we care.. and because you are killing our budget.

Token Efficiency And Withdrawal Symptoms

Leaving aside the dystopian, yet very real, discussion about replacing human salaries with token budgets, we should ask a more practical question: How do tech organizations become more token efficient now that we are all already so deeply hooked on the product, without everything falling apart?

Unfortunately, the deck is stacked against us for several reasons:

The harness is owned by the token peddlers

Most companies are using the harnesses provided by the same AI companies selling the tokens. But the harness is also the biggest driver to the compounding effect of token usage, which means we have an inherent conflict of interest. So token efficiency not only becomes harder to implement, it also becomes more opaque and harder to analyze almost by design, and there’s no incentive on the other end to solve this issue.

We over democratized the agentic toolbox

As with all things fueled by FOMO, governance and standardization have mostly been an afterthought in deploying these agents and ecosystem around them. What this means is that every person in your company using Codex or Claude today has a different mixture of AGENT.md, skills, MCP servers. Many times overlapping, many times vibe coded themselves to immense sizes in context usage. With the end user usually having very little awareness to context bloat and related costs. And don’t even get me started on discretionary model choice, and the fact people always go to the top shelf for even the most pedestrian needs.

The models cover for your incompetence and also for their own with token usage

Frontier model providers pride themselves on the time horizon of their models. That is to say, they can increasingly perform tasks that would have taken more time for a human expert to perform them. Much less celebrated are benchmarks for cost-per-outcome and token usage, and very few of them that reflect reality. In reality there’s a lot less “one shotting” than you’d think and a lot more ambiguity. The agentic harnesses quite often end up in a token spending spiral in an attempt to solve a problem it lacks input to be able to solve because you neglected to provide that input, or simply it can’t solve it with its current reasoning capabilities.

Tokens became part of the expected working conditions for engineers

Engineers are a species that gets used to the good stuff very fast and are very hard to let go when it comes to quality of life. Simply put, if your company is not willing to pay for their performance enhancing tokens, they will go to another company that will. So on top of everything, cutting back on token spend can mean not only productivity risk, but also retention risk for top talent.

So the risks of the Tokenpocalypse are very real but also the headwinds for mitigating them, and assuming tokens are not going to magically become 10 times cheaper in the next 6 months, there’s hard work we need to do very fast so this doesn’t end up in some form of a catastrophe

So What Now?

The solution is definitely not telling engineers to simply stop using agents, or even asking them to self govern and use less tokens. That train has long left the station.

I’m not going to tell you that I have all the answers, or even most of them. I can only assume that as you are reading this, there are already several dozen early stage startups sharpening their investor deck that explains how they plan on solving this exact problem. But until those come to save us, there are a few things that look like safe bets:

Treat tokens like infrastructure, not like office supplies

The first step is simply changing how organizations account for token spending. Many engineering organizations are currently using multiple types of harnesses and model providers, which is the smart thing to do, because you definitely don’t want to have all your eggs in one basket. However that makes token cost tracking and harness tracking a distributed problem. Companies need to start investing in collecting this usage data and track it in a unified way so that token usage is visible, attributed and owned. Owned by a team, org function or use case.

You can’t optimize what you can’t measure. The same FinOps discipline we eventually applied to cloud compute sprawl and the big data explosion needs to be applied here, and probably faster.

Build a governance layer before the chaos compounds

Secondly, and this one is much harder in practice, we need to identify systematic misuse. Identify users that use Opus regardless of the task at hand, and start sessions with their context window already 50% full with crap. Identify bad skills that create context bloat or workflow overhead, identify bad MCP servers with ridiculous metadata and LLM unfriendly response payloads.

You can use centralized infrastructure such as MCP proxies and toolkit servers, as well as AI plugin managers, to give yourself both better visibility and a more sane way to centrally control context bloat by over tooling. But to be frank this area still leaves more to be desired.

Promote a token meritocracy culture

As “token budget” becomes an acceptable term to use in the office space, so should the understanding of how that budget is allocated. We live in a time of extreme asymmetry to how different people are able to leverage these tokens to create value, and therefore the way we allocate budget should also be asymmetrical. But that asymmetry should be led by outcomes and value created, rather than the circular logic of tokenmaxxing where more tokens used means more tokens given.

Simply put, those that are able to bring more value with these tokens, are the ones to be allowed more of the token pie. The rest will need to learn to optimize.

Explore the open source and open weights ecosystem

Yes, both the best models right now and the best harnesses are owned by the frontier model labs. But the gaps are fast being closed both by open weights models that cost a fraction, and OSS harnesses like OpenCode, that some already swear to be as good if not better than Claude.

I expect an explosion in this area in the coming months, as the entire industry is scrambling for alternative solutions, as well as innovating in places where the frontier labs have no motivation to invest in like better optimizations, observability and smart routing of tasks across models.

Not all fixes are that easy

Not all fixes are that easy

None of this is a silver bullet. And none of it addresses the deeper, stranger question lurking underneath all of this: what happens to an industry whose cognitive output has become genuinely dependent on an input it can’t fully control or predict the cost of?

As for me, I’m still drinking the coffee. I’m still burning the tokens. Some habits are worth the price.

The price of coffee fluctuates, but it definitely doesn’t compound recursively. It doesn’t spawn sub-agents that make themselves coffee. And the global coffee supply definitely isn’t controlled by a handful of companies that also have their hand on the scale of how much coffee I end up consuming.

But my employer is still paying for my coffee, and still paying for my tokens, so as far as I’m concerned everything is fine. For now.


메타데이터
post_id
53e87a7daefe
slug
the-tokenpocalypse-is-here-53e87a7daefe
url
https://medium.com/wix-engineering/the-tokenpocalypse-is-here-53e87a7daefe
canonical_url
https://medium.com/wix-engineering/the-tokenpocalypse-is-here-53e87a7daefe
author_url
https://medium.com/@yossigoldstein
status
ok
fetched_at
2026-07-09 15:12:33