Claude Sonnet 5 Is Here: I Just Tested It (And It’s Closer to Opus Than I Expected)
Anthropic has just released Claude Sonnet 5 as the smartest model for everyday work. It’s available in Claude Code and as the default…
Claude Sonnet 5 Is Here: I Just Tested It (And It’s Closer to Opus Than I Expected)

Anthropic has just released Claude Sonnet 5 as the smartest model for everyday work. It’s available in Claude Code and as the default model across platforms.
I went through the Claude Sonnet 5 system card the moment it was released to see how different it is from Claude Sonnet 4.6, which is my daily use model.
If you are not a premium Medium member, read the full story here for FREE , but consider joining Medium to support my work — Thank you!
Sonnet 5 closes the gap with Opus 4.8 in a way we have not seen in the previous Claude Sonnet models.

- On agentic coding, it scores 63.2%, compared to Opus 4.8’s 69.2%
- On knowledge work, it edges past Opus 4.8
That is a big claim for a model priced at $2 per million input tokens and $10 per million output tokens through August 31, less than half of what Opus 4.8 costs.
It is live everywhere today, and I just tested my favourite spot, Claude Code, and immediately, I launched the updated version. I noticed this message:

It's also the default model for Free and Pro plans, available on Max, Team, and Enterprise, and already in the Claude Platform.

It’s available in the free version as well, as you can see in the image above, and it's set as the default model.
I switched to it immediately on Claude Code and ran it through testing.

Let me break down what’s new and what I learned from the Claude Sonnet 5 system card.
Claude Sonnet 5
I started my chat test with the popular question :
How many rs in strawberry

I followed up with the misspelled version of the question :
How many rs in strawberrry

And you can see the results.
Here are the improvements in the new Claude Sonnet 5 from 4.6.
1) Most Agentic Sonnet Yet
Anthropic calls Sonnet 5 its most agentic Sonnet model.
It can make plans, use tools like browsers and terminals, and run autonomously at a level that used to require Opus class models.
The agentic era started with Sonnet; Claude Sonnet 3.5, 3.6, and 3.7 were the first models that showed skill in coding and tool use.
The last few improvement in agentic capability went to Opus, but now Sonnet 5 has closed that gap.
2) Close Benchmark Gap With Opus 4.8
As I mentioned in the intro above, Claude Sonnet 5 on agentic coding scores 63.2%, compared to Opus 4.8 at 69.2% and Sonnet 4.6 at 58.1%.
On knowledge work, Sonnet 5 passes Opus 4.8

Opus 4.8 still leads on raw accuracy across most tasks, but the difference between the two models is the smallest it has ever been between a Sonnet and an Opus release.
3) Introductory Pricing
Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens through August 31, 2026.
After that, it moves to standard pricing at $3 input and $15 output.
For comparison, Opus 4.8 runs at $5 input and $25 output.
Even at standard pricing after August, Sonnet 5 costs 60% less than Opus 4.8.
4) Default Model for Free and Pro
Sonnet 5 is now the default model for anyone on Free or Pro plans.
Max, Team, and Enterprise users also have full access. It is live in Claude Code and on the Claude Platform starting today.
Developers can call it through the API with claude-sonnet-5
5) Finishes What It Starts
Early access partners reported Sonnet 5 finishes complex tasks where previous Sonnet models would stop.
Daniel Shepard, a senior engineer at Zapier, described handing Sonnet 5 a two-part job: update Salesforce account tiers, then send a launch announcement to enterprise contacts. It finished end-to-end.
6) Self Checks Without Being Asked
Testers also noticed Sonnet 5 checks its own output without being instructed.
One example from Anthropic’s partners involved Sonnet 5 investigating a bug. Unprompted, it wrote a reproducing test, implemented the fix, then stashed the fix to confirm the bug came back without it.
7) Lower Rate of Undesirable Behaviors
The safety assessments show Sonnet 5 has an overall lower rate of undesirable behaviors than Sonnet 4.6.
It is better at refusing malicious requests and resisting hijack attempts in prompt injection attacks. Hallucination and sycophancy rates are both down compared to Sonnet 4.6.

It still shows somewhat higher rates of misaligned behavior than Opus 4.8 and Mythos Preview on Anthropic’s automated behavioral audit, but lower than its predecessor.
8) Cyber Safeguards Enabled by Default
Sonnet 5 was not trained on cybersecurity tasks, but its general intelligence gains carried over into cybersecurity.
For example, on the Firefox 147 exploit evaluation built with Mozilla, Sonnet 5 never produced a full working exploit, the same as Sonnet 4.6, but showed a higher partial success rate.
Anthropic launched Sonnet 5 with the same real-time cyber safeguards used in Opus 4.7 and 4.8, which are less strict than those on Fable 5, since the risk level is considered low.
9) Higher Rate Limits
Anthropic increased rate limits across Chat, Cowork, Claude Code, and the Claude Platform to accommodate the higher token usage that comes with running Sonnet 5 at higher effort levels.
If you push Sonnet 5 to Extra High effort, you are going to burn more tokens per task, and Anthropic has just adjusted the limits.
10) Cost Performance Curve
Anthropic published cost performance charts comparing Sonnet 5, Sonnet 4.6, and Opus 4.8 across different effort levels on agentic search and computer use evaluations.

At its Extra High effort setting, Sonnet 5 performs in line with Opus 4.8 at medium to high effort.
Opus 4.8 still wins on raw accuracy at its own higher settings, but Sonnet 5 at xhigh covers ground that used to be only an Opus zone.
Final Thoughts
Claude Sonnet 5 is the best Sonnet release Anthropic has shipped, and the pricing makes it affordable.
The benchmark gap with Opus has never been this close on a Sonnet model, and early testers' results align with the benchmarks.
I am currently running Sonnet 4.6 as my daily model, and have just switched to Sonnet 5 to see its value proposition.
I expect that Sonnet 4.6 will soon be phased out and replaced with Sonnet 5, as was the case with Sonnet 4.5.
I have also just started testing Sonnet 5 on Claude Code and will share more details in the next test article.
Have you tested Claude Sonnet 5 yet? Let me know in the comments.
메타데이터
- post_id
- 3fe2a9a0c415
- slug
- claude-sonnet-5-is-here-i-just-tested-it-and-its-closer-to-opus-than-i-expected-3fe2a9a0c415
- url
- https://medium.com/ai-software-engineer/claude-sonnet-5-is-here-i-just-tested-it-and-its-closer-to-opus-than-i-expected-3fe2a9a0c415
- canonical_url
- https://medium.com/ai-software-engineer/claude-sonnet-5-is-here-i-just-tested-it-and-its-closer-to-opus-than-i-expected-3fe2a9a0c415
- author_url
- https://medium.com/@joe.njenga
- status
- ok
- fetched_at
- 2026-07-09 10:05:04