← Back to list

Why Old AI Models Suddenly Feel Stupid.

The strange psychology of moving baselines, model drift, and our shrinking patience for yesterday’s magic.

Carson · 2026-07-18 14:05 · 0 claps · 10.0 min read
#ai-model #machine-learning #technology-trends #workflow-automation #writing-tools
Open on Medium ↗
Wiki topics: ML · Machine Learning PSY · Psychology EDU · Education & Learning

Why Old AI Models Suddenly Feel Stupid.

The strange psychology of moving baselines, model drift, and our shrinking patience for yesterday’s magic.


👋 Hey, I’m Ian. A Staff engineer running my startup with agents. I write about practical AI Agents that make money — AI agents, autonomous workflows, and what happens when companies start putting agents on payroll.

I’m also building **Narrareach**, a distribution platform for writers and founders. It lets you schedule Substack Notes and articles, cross-post to all social media platforms, publish across Medium, Linkedin newsletters and Substack, and see what’s actually working so you can double down on the right ideas.

Batch schedule your notes and articles then cross-post everywhere:

P.S. If you’re building an agent, hiring AI into a workflow, or trying to grow your writing without manually posting everywhere, reply to this email. I may feature a few interesting examples in future deep dives.


A strange thing happens after you use a better AI model for a few days.

You go back to the old one, ask it something normal, and suddenly it feels like it has been hit on the head with a frying pan.

This is confusing because, on paper, nothing dramatic may have happened. The old model still has the same name. The button still looks the same. The company has not sent you a polite email saying, “We regret to inform you that your favorite model now has the reasoning ability of a sleepy intern.” Everything appears stable.

Then you ask it to write code, explain a bug, summarize a document, or help you think through a decision, and the answer feels off. Slower. Flatter. Less aware. It misses the thing the new model would have caught immediately. It gives you the sort of response that makes you stare at the screen and wonder if you accidentally opened the diet version.

This feeling is everywhere now. A new flagship model drops, people try it, and within a week the previous model starts getting treated like old furniture. The same model that felt brilliant last month suddenly feels small. The old genius becomes the new intern.

The obvious question is whether the old model actually got worse.

The more interesting question is why it feels that way even when it didn’t.

Retro robot questions its relevance in AI evolution.

Retro robot questions its relevance in AI evolution.


Yesterday’s magic becomes today’s minimum standard

Human beings are terrible at staying impressed.

The first time a model writes a clean paragraph, you feel like you are watching electricity discover literature. A few weeks later, you get annoyed because it used the word “delve” again and somehow turned a simple idea into a LinkedIn leadership retreat.

That is not because the magic disappeared. It is because your baseline moved.

Every powerful tool trains your expectations upward. Once you experience a model that catches subtle bugs, understands your tone, follows a long argument, and holds context without wandering into a nearby bush, you do not return to the older model as the same person. You return with new standards.

The old model may not have changed.

You did.

This is the cruel little joke of progress. The better tool does not only improve the future. It also edits your memory of the past. It makes yesterday’s miracle look clumsy by comparison.

This happens in every category. The phone camera that once looked incredible now produces photos that feel like they were taken through a sandwich bag. The laptop that once felt fast now sounds like it is preparing for takeoff when you open three tabs. The car that once felt luxurious now lacks features you suddenly consider basic, even though you survived thirty years without heated steering wheels and did not die once.

AI is no different. We are just moving through the cycle faster because the gap between models is not measured in decades. It is measured in months, sometimes weeks, sometimes one product launch that ruins your entire relationship with the model you used to defend online.


The model may also be changing

There is another possibility, and we should not pretend it is imaginary.

Sometimes the model really does change.

The product you use is not a frozen marble statue sitting quietly in a museum. It is usually part of a living service. Behind the same model name, companies may change routing, safety behavior, context handling, system prompts, tool access, latency tradeoffs, memory behavior, inference settings, or infrastructure priorities.

To the user, all of this appears as one simple thing: “Why is it worse?”

That is the frustrating part. You rarely get a clean explanation. You do not see a dashboard that says, “Today’s answer quality is 12% flatter because we adjusted serving behavior while trying to reduce latency.” You see the same logo, the same model name, and a response that makes you want to take a walk.

This is why people are not crazy when they say a model feels different. Commercial AI systems can drift. API behavior can shift. Safety tuning can alter tone. Product changes can make the same named model behave differently from week to week. Even when the core weights stay the same, the experience around the model can move.

That distinction matters.

When users say, “The model got dumb,” they may be describing several different things at once. The model may be weaker at a task. The system prompt may have changed. The routing may be different. The context window may be handled differently. A safety layer may be more aggressive. The user may be asking harder questions. Their standard may have risen.

From the outside, all of these feel almost identical.

The answer disappoints you.

That is the whole crime scene.


We judge intelligence comparatively

Intelligence is not experienced in isolation. We judge it against whatever we just saw.

A high schooler sounds brilliant to a child and confused to a professor. A junior engineer sounds sharp until you watch a staff engineer solve the same problem in seven minutes while eating a yogurt and saying almost nothing. A model that felt smart beside yesterday’s tools can feel painfully average beside the new one.

This does not mean the old model is useless. It means comparison changed the room.

AI makes this especially dramatic because the difference between models often appears in small places. The new model notices the implied constraint. It catches the hidden contradiction. It remembers the thread of your argument. It does not just answer the question; it understands the shape of the work around the question.

Then you return to the older model and feel every missing layer.

It answers, but it does not quite see.

That is often the difference people are reacting to. The older model can still produce something plausible, but the newer one has trained you to expect a kind of situational awareness. Once you get used to that, anything less feels like talking to someone who technically heard you but was also checking their phone.


The task changed too

There is one more reason old models suddenly feel worse.

We promote them into harder work.

A model may have felt excellent when you were asking for summaries, drafts, or simple explanations. Then a new model arrives and shows you that AI can help with deeper tasks: multi-step coding, debugging, research synthesis, planning, architectural decisions, workflow automation, or agentic work that runs longer than one cute chat session.

After that, you start asking every model to operate at that higher level.

The old model did not necessarily decline. You gave it a harder job.

This happens to people too. Someone can be great at one level and visibly strained at the next. A solid individual contributor may struggle as a manager. A strong manager may struggle as an executive. A brilliant writer of short posts may collapse when asked to write a book. Capability is often contextual.

AI models are the same. They have zones where they feel sharp and zones where their weaknesses become impossible to unsee. New models expand the zone. Older models get dragged into the new expectations and blamed for not surviving the promotion.

That is why “same model, same weights” does not settle the question. The model can stay the same while the job changes around it.


Better models expose what we were tolerating

The most painful part of progress is that it reveals our previous compromises.

Before you use a stronger model, you tolerate more. You rewrite more. You explain more. You accept shallow reasoning because it is still better than starting from a blank page. You forgive the missed nuance because at least it gave you momentum.

Then a better model arrives and suddenly you realize how much unpaid management you were doing.

You were carrying the context.

You were repairing the reasoning.

You were translating your own instructions into something the model could follow.

You were cleaning up the tone.

You were catching the errors.

You were the safety net, editor, project manager, and emotional support animal for a machine with confidence issues.

Once a newer model removes even part of that burden, the old workflow starts to feel exhausting. The older model is not merely worse in output. It demands more supervision. It creates more cleanup. It makes you do more invisible labor.

That is why people get irritated so quickly. They are not only judging intelligence. They are judging how much work the model leaves behind.

A weaker answer is annoying.

A weaker answer that creates more work is betrayal with a subscription fee.


The trust problem

This is where AI companies have a real problem.

Users can accept that models change. They can accept that new systems are better. They can accept that some older models will become cheaper, slower, deprecated, or moved into a different tier.

What they struggle with is uncertainty.

If a model changes and nobody tells you, you start inventing explanations. Maybe it was nerfed. Maybe the company reduced compute. Maybe routing changed. Maybe safety got heavier. Maybe your prompts are worse. Maybe you have become spoiled. Maybe Mercury is in retrograde and the datacenter has vibes.

The lack of clarity turns ordinary product evolution into folklore.

This is especially damaging because people now build real workflows around these models. Developers rely on them for coding. Writers rely on them for drafting and editing. Founders rely on them for research and operations. Teams build agents, automations, and evaluation loops around specific behaviors.

When the behavior changes, trust changes with it.

A model does not need to collapse for users to lose confidence. It only needs to become unpredictable in the exact places they learned to depend on it.

That is why AI products need better versioning, clearer change logs, and more honest language around model behavior. Not every user needs a research paper. They just need to know whether the thing they are paying for changed in a way that could affect their work.

“Same name” is no longer enough.


The spoiled user is also real

Still, we should be honest with ourselves.

Sometimes we are spoiled.

A model that would have looked impossible three years ago now gets insulted because it takes a few extra seconds to understand a messy prompt we wrote while eating lunch. We ask for strategy, code review, literary editing, market analysis, therapy-adjacent emotional support, and a meal plan, then complain that it missed the tone slightly.

There is something absurd about this.

We are becoming impatient with miracles because the miracles are improving too quickly.

That does not mean the complaints are wrong. Many are legitimate. It does mean our emotional relationship with AI is getting strange. We are starting to treat intelligence like loading speed. If it does not arrive instantly, in the right format, with the right amount of nuance, we act like civilization has failed.

The user is not always the victim in this story.

Sometimes the model is fine.

Sometimes we are standing in front of a machine that can read, write, code, reason, summarize, translate, and explain almost anything, muttering, “Why is this thing so dumb today?”

Progress has made us rude.


How to think about it

The useful move is to stop asking one blunt question.

“Did the model get worse?”

That question is too small.

A better question is: what changed?

Did the model change?

Did the product wrapper change?

Did routing or safety behavior change?

Did the task become harder?

Did my standards rise?

Did the newer model teach me to notice flaws I used to forgive?

That mental model is more helpful than panic. It also keeps you from becoming the person who treats every weak answer as proof of a grand corporate conspiracy. Sometimes the provider changed something. Sometimes your prompt is bad. Sometimes the old model is still good, just not for the job you are now trying to make it do.

The practical answer is to test models against your own work. Keep a small set of prompts that matter to you: one coding task, one writing task, one reasoning task, one summarization task, one weird edge case that usually exposes weakness. Run them when a new model drops. Run them again when an old model starts feeling off.

Do not rely only on vibes.

Vibes are useful. Vibes tell you where to look. But if your entire evaluation system is “it feels lobotomized today,” you are one bad afternoon away from becoming a Reddit thread with a pulse.


Yesterday’s genius, today’s intern

Old AI models suddenly feel stupid because intelligence is comparative, product systems change, and our patience shrinks every time the frontier moves.

All three can be true at once.

The old model may have drifted.

The product may have changed.

Your standards may have moved.

Your task may have become harder.

Your taste may have improved.

That is the strange psychological tax of rapid progress. Every new model does more than raise the ceiling. It lowers our tolerance for the floor. It rewrites the emotional memory of what came before.

Yesterday’s magic becomes today’s minimum standard.

Yesterday’s genius becomes today’s intern.

And the most unsettling part is that we will do the same thing to today’s best model soon enough.

We will praise it, depend on it, build workflows around it, recommend it to friends, defend it in comment sections, and then one day a better model will arrive.

Within a week, we will go back to the old one and wonder what happened.

Nothing mystical.

Nothing necessarily malicious.

But progress doing what progress always does.

It gives us a better tool, then makes us impatient with the one that taught us to believe.


P.S. I like using Substack Notes to test smaller ideas before turning them into full articles, but writing one Note at a time can quickly turn into another tab I have to live inside.

That is where Narrareach has been useful. It gives writers a way to batch schedule Substack Notes, import Substack Notes from CSV, and manage a Substack Notes content calendar without manually sitting inside the Substack composer every day.

Native Substack scheduling is useful for one-off Notes. Narrareach is built for the heavier workflow: bulk schedule Substack Notes, repurpose articles into multiple follow-up Notes, choose a best-time-to-post cadence, cross-post the strongest ideas, and track Substack Notes analytics so you can see which posts actually bring subscribers back to the publication.


메타데이터
post_id
50535d47bea7
slug
why-old-ai-models-suddenly-feel-stupid-50535d47bea7
url
https://medium.com/@kiprono.ca001/why-old-ai-models-suddenly-feel-stupid-50535d47bea7
canonical_url
https://medium.com/@kiprono.ca001/why-old-ai-models-suddenly-feel-stupid-50535d47bea7
author_url
https://medium.com/@kiprono.ca001
status
ok
fetched_at
2026-07-25 12:44:45