← Back to list

The Best Model Is the One You Know How to Use

Hot take:,Benchmarks do not matter to 99% of us. LinkedIn driven AI Brofluence is hurting your bottom line.

Marco Kotrotsos in Autocomplete. Real World AI · 2026-07-10 15:56 · 83 claps · 6.1 min read paywalled
#ai #artificial-intelligence #productivity #programming #software-development
Open on Medium ↗
Wiki topics: EVAL · Evaluation & Benchmarks AI · AI · General 💻 · Programming ⏱️ · Productivity

The Best Model Is the One You Know How to Use

Hot take:,Benchmarks do not matter to 99% of us. LinkedIn driven AI Brofluence is hurting your bottom line.

There is a genre of content that runs on a two-week loop. A new model drops, tops a benchmark by a few points, and the takes begin. This one is the new king. That one is finished. Switch now or fall behind. The old favorite is dead. Then, roughly a fortnight later, a different lab ships something, the numbers flip, and the exact same people write the exact same posts with the names swapped.

I want to say the quiet thing plainly. For almost everyone actually using these tools to get work done, none of it matters. The percentage points do nothing for you. The leaderboard reshuffling is sports commentary for people who do not play the sport.

-none of it matters

The move that actually compounds is boring and unglamorous: pick a capable model, and get genuinely good at using it. Learn its habits, its failure modes, the way it likes to be asked. That is worth more than every leaderboard swap you will be tempted by this year, and it is not close.

Every now and then the move is significant. Like Opus to Fable. But even then- give it a fortnight and your favourite lab catches up.

Three things to take away

  • The benchmark gap between the top models is smaller than the gap your own skill creates. How well you prompt and steer a model moves your results more than the difference between the number one and number three spot. You are the bigger variable.
  • Every switch resets the intuition you were building. Model fluency is a compounding asset. Hopping to chase a new leader throws that progress away and starts you back at zero, over and over.
  • Any frontier model is good enough to go deep on. Claude, GPT, Gemini, a strong open model, the choice matters far less than the depth. Pick one you like and stop shopping.

What the benchmark points actually measure

Look at what tops the charts and you will notice the tests are things like graduate-level reasoning and abstract math. Impressive, and almost entirely disconnected from what you do on a Tuesday, which is more like drafting a document, refactoring a function, summarizing a thread, or talking through a decision. A model can gain two points on a math benchmark and change nothing about how it handles your actual work.

It gets worse for the leaderboard’s credibility. A 2025 study politely titled “The Leaderboard Illusion” showed that models can climb the rankings without genuinely getting better, because the rankings can be gamed. The number went up. The thing you care about did not.

And the part that should end the argument is this. Research on prompt sensitivity keeps finding that the same model produces meaningfully different results depending on how it is asked, enough that small changes in phrasing can reshuffle where models land relative to each other. Sit with that. The variation created by how you prompt can be larger than the gap between the models you are agonizing over. Which means the highest-leverage thing in the whole system is not the model you pick. It is how well you have learned to drive it.

The benchmark gap between the top model and the number three model is a rounding error, while the gap between a casual and a fluent user of the same model is the one that decides your output

What hopping actually costs you

Switching models feels free. It is not. It just bills you in a currency you are not counting.

When you have used one model seriously for a while, you have built a mental model of it. You know where it tends to overreach and where it quietly gets lazy. You know it pads its answers a certain way, or that it needs the constraint stated twice, or that it is great at this kind of task and unreliable at that one. You have a folder of prompts that work. You have a setup around it, the harness, the shortcuts, the little rituals, that fits your hand.

Every one of those is model-specific, and every one of them resets when you jump. The new model has different failure modes you have not mapped yet, a different verbosity, a different idea of what you meant. Your prompts need re-tuning. Your intuition, the thing that lets you glance at an output and instantly know whether to trust it, is gone, and you are back to checking everything carefully because you do not yet have the feel. You traded a compounding asset for a shiny number, and you will do it again in two weeks when the number moves.

That is the trap. The churn guarantees you never get past the beginner phase with anything, because you keep restarting the clock.

What fluency looks like, and why it is the real edge

Get good at one model and something changes in how you work. You stop fighting it. You know the shape of a request it will nail on the first try versus one you need to break into steps. You know when its confidence is earned and when it is bluffing. You develop the reflex of steering, of nudging it back on track with a half-sentence instead of a paragraph. Your prompts get shorter and your results get better at the same time, which only happens when you actually understand the thing.

This is where the real productivity lives. Not in having this quarter’s top-ranked model, but in being the person who can get 90 percent of a frontier model’s capability out of it reliably, because most people are leaving half of it on the table by never learning it properly. The gap between a fluent user and a casual one, on the same model, is enormous. The gap between the top two models, for your work, is usually a rounding error.

Depth beats novelty. It is not even a close contest, and the churn exists specifically to make you forget that.

What this does not mean

This is not brand loyalty, and it is not “never switch.” That would be its own kind of foolish. There are real reasons to move.

Switch when there is a genuine capability gap that blocks your actual work, the model simply cannot do a thing you need, not the model scored lower on a test. Switch when the economics change enough to matter for how much you use it. Switch when a new class of feature, not a new decimal place, opens up a workflow you could not do before. Those are decisions grounded in your work, and you will feel them as friction you actually hit, not as a headline you read.

What you should not do is switch because a chart moved, because a thread declared your model dead, or because the vibe online shifted this week. Evaluate on your own tasks, not on someone else’s benchmark. If you must compare, compare by running your real work through the candidates and judging the outputs yourself. That is the only leaderboard that describes your life.

And to be clear about the premise: the model you go deep on can be almost anything current. This is not an argument for a specific lab. Anthropic, OpenAI, Google, a capable open model, all of them are more than good enough to build serious fluency on. The winning move is choosing one and committing, not which one you choose.

You are the bigger lever. Depth beats the leaderboard everytime.

The scoreboard is not the game

The people writing “this model killed that model” every other week are, mostly, in the engagement business. That is content, not information. It is optimized to make you feel behind, because feeling behind is what makes you click, and click again when the ranking flips back.

You do not have to play. Pick a model you like working with. Spend the hours you would have spent reading model-versus-model takes actually using it, pushing it, learning where it bends. In a month you will be getting more out of it than the people who switched three times in the same period and never got past the tutorial with any of them.

In two weeks something new will top the charts, and someone will tell you your choice is finished. It will not be. You will still be fluent, and fluency is the thing that was ever going to move the needle.

Marco Kotrotsos, specializing in practical AI implementation for organizations ready to close the gap between AI hype and AI value. With 30 years of IT experience now focused purely on AI deployment, he works hands-on with companies to turn AI potential into measurable business outcomes.

This article is published in Autocomplete, a Medium publication about real-world AI for practitioners and decision-makers. We’re always looking for writers. If you’re building with AI and have something worth sharing, reach out.

My free Substack newsletter, also called Autocomplete, can be found here: https://acdigest.substack.com.

My books on Amazon: Claude Code for Everyone Else and From Vibe to Production.

I also take on a small number of mentees one-on-one on MentorCruise.


메타데이터
post_id
bdfb88e4130a
slug
the-best-model-is-the-one-you-know-how-to-use-bdfb88e4130a
url
https://medium.com/autocomplete-real-world-ai/the-best-model-is-the-one-you-know-how-to-use-bdfb88e4130a
canonical_url
https://medium.com/autocomplete-real-world-ai/the-best-model-is-the-one-you-know-how-to-use-bdfb88e4130a
author_url
https://medium.com/@kotrotsos
status
ok
fetched_at
2026-07-13 06:23:13