People Are Learning to Talk to AI Like Cavemen to Save Tokens
A Reddit trick for stretching your rate limit, and what it says about who’s adapting to whom.
People Are Learning to Talk to AI Like Cavemen to Save Tokens
A Reddit trick for stretching your rate limit, and what it says about who’s adapting to whom.
The trick went around Reddit a few months back. Strip your prompts down to fragments. No pleasantries, no filler, no full sentences. Tell the model to answer the same way — short, blunt, no narration. People tried it, found it genuinely stretched their token budget further before hitting a rate limit, and the whole thing turned into a meme. Caveman-speak as a survival strategy for talking to a multi-billion-dollar AI product.
It’s a funny trick. It’s also a small, clean example of something worth noticing.

Same person. Same intelligence. Fewer words, because the meter’s running.
What actually happened here
Somewhere in the last year, the economics of running these models stopped being invisible to the people using them. Rate limits got tighter and more visible. Anthropic pushed heavy users toward separately billed “extra usage.” Anthropic also drew a harder line around third-party agent harnesses — tools like OpenClaw that keep a model running for long, autonomous stretches — after describing the load they put on its systems as an “outsized strain,” pushing that kind of usage off the standard subscription entirely and onto metered billing instead.
Inference costs money, infrastructure has limits, and a company deciding where to draw the line between “included” and “billed separately” is running a normal business. What’s interesting is what happened on the other side of that line. Users didn’t push back on the constraint. They adapted their own behavior to fit inside it. They changed how they type.
A cost structure that lives entirely on a vendor’s servers, invisible to any individual user, reached backward through the interface and quietly reshaped how people communicate with it. Not through a policy change or an announcement — just through friction, absorbed calmly enough that the response to it became a meme rather than a complaint.
The advice everyone’s writing right now
Look at how the current advice is framed. Article after article walks through the mechanics — context window management, when to spin up a subagent, how caching works, what a “chatty” agent costs you in output tokens. All of it is accurate, some of it genuinely useful. Almost none of it pauses on the prior question: why is the user the one doing this work at all?
A rate limit is a constraint the vendor chose, sized around infrastructure the vendor controls, priced by a formula the vendor sets and can change without much notice. The instinct across the entire discourse — mine included, until I sat with this for a while — was to treat that constraint as weather. Something to route around, optimize against, get clever with. Nobody’s first move was to ask whether the constraint itself was drawn in the right place, because the constraint doesn’t feel like a decision. It feels like physics.
It isn’t physics. It’s a business model wearing physics as a costume. And the giveaway is that the “fix” being passed around isn’t a better product decision from the vendor — it’s users learning to write worse prose so a meter runs slower.
Why this is bigger than one Reddit meme
The caveman-speak trick is trivial on its own. Strip filler words, save a few thousand tokens, get a slightly longer session before the wall. Fine. But it’s a clean, almost comic instance of a much larger pattern that’s easy to miss because it usually doesn’t come with a funny quote attached.
Structured JSON output over free-form text, because it’s “more efficient,” even when free-form would communicate the actual nuance better. Dynamic turn limits that cut an agent off mid-reasoning because a Stevens Institute study showed cost drops 24% that way, quality be damned in the margin. Teams designing entire agent architectures around minimizing calls to a model, not because fewer calls produce a better outcome, but because each call has a price tag attached to someone else’s balance sheet. Somewhere along the way, “how do we get the right answer” quietly became “how do we get an acceptable answer inside someone else’s budget.”
That’s not a criticism of any single team making a single sensible tradeoff. Cost discipline is real and it matters. The point is narrower: an entire category of technical decision-making is now downstream of a pricing model none of the people making those decisions had any say in setting. And the industry’s response to noticing this has mostly been to get better at compliance. Better prompt compression. Better caching. Better turn limits. All of it aimed at fitting more comfortably inside a box, not at asking who built the box or whether it’s the right size.

The marketing number and the tested number aren’t measuring the same trick.
A smaller ask than it sounds
None of this is an argument against rate limits, or against the idea that inference costs real money and companies need to manage that. Compute isn’t free, frontier models cost more, and a business that doesn’t account for that carefully doesn’t stay a business for long.
The smaller thing worth asking for is just visibility into the tradeoff. If trimming a prompt to fragments buys a longer session, that’s a useful fact on its own — and it’s also fine to notice, without much drama, that the fix here is on the user’s side of the interface rather than the product’s. If a turn limit saves cost at some measurable expense to output quality, that’s a reasonable design choice, and probably one worth stating plainly rather than discovering mid-task.
The caveman-speak trend will fade the way memes do. The underlying pattern — people quietly adjusting their own behavior around a cost structure they can’t see — will likely keep showing up in new forms, mostly unremarked on, because it rarely comes wrapped in something as shareable as a joke about grunting at a chatbot.
I write about building real AI products and the thinking behind the decisions. If this is useful, follow along on Medium and Substack ✌️
메타데이터
- post_id
- 0db177f0a323
- slug
- people-are-learning-to-talk-to-ai-like-cavemen-to-save-tokens-0db177f0a323
- url
- https://medium.com/@bhavyansh001/people-are-learning-to-talk-to-ai-like-cavemen-to-save-tokens-0db177f0a323
- canonical_url
- https://medium.com/@bhavyansh001/people-are-learning-to-talk-to-ai-like-cavemen-to-save-tokens-0db177f0a323
- author_url
- https://medium.com/@bhavyansh001
- status
- ok
- fetched_at
- 2026-07-20 00:44:07