When all recommending Claude then this “Sakana AI” Happened.
A Tokyo startup just matched Anthropic’s best models without ever touching them. I’m still sitting with that.
When all recommending Claude then this “Sakana AI” Happened.
A Tokyo startup just matched Anthropic’s best models without ever touching them. I’m still sitting with that.

I have a default answer when someone ask me which AI model to build on.
I’ve had it for couple of years. It comes out automatically now — the reasoning, the tradeoffs, the “here’s why this one for this use case.” I’ve built pipelines, automations, entire client workflows around it. My credibility, in part, sits on the quality of that recommendation.
Last week a Japanese startup called Sakana AI launched something called Fugu. And for the first time in a while I found myself genuinely unsure whether my default answer is still right.
Not panicking. Just… unsure. Which, honestly, is its own kind of uncomfortable.
Here’s what Sakana did that I can’t stop thinking about.
Anthropic’s two most powerful models right now Fable 5 and Mythos Preview — are not publicly accessible. Export controls. A significant chunk of the global developer community simply cannot use them. I’ve had clients ask about them. The answer has been: not available to you.
Fugu Ultra matched them on benchmarks. Without using them. Without having access to them.
It didn’t do this by training a bigger model. It did it by building a system that orchestrates the best available models like GPT, Claude Opus, Gemini dynamically, at inference time, assigning them roles: one thinks, one executes, one verifies. The coordination itself is learned, not hand-designed.
The result comes back through one API as if it came from one model.
And on SWE-Bench Pro — one of the harder engineering benchmarks Fugu Ultra scored 73.7%. Claude Opus 4.8 scored 69.2%. GPT-5.5 scored 58.6%.
I want to be careful here because benchmarks are benchmarks and real workloads are different within 24 hours of launch, independent testers found gaps between Sakana’s claims and actual use.
So I’m not saying Fugu is better than Claude.
I’m saying a company in Tokyo, founded three years ago, just posted numbers that made me look twice at a recommendation I’ve been giving on autopilot.
That’s the part that got me.
The founders of Sakana AI (Llion Jones) is a co-author of the original “Attention Is All You Need” paper, and David Ha, former head of research at Stability AI that left large AI companies specifically because they were frustrated with the industry’s obsession with scaling single massive models.
Their thesis from day one was the most powerful AI systems won’t be isolated monoliths. They’ll be collaborative ecosystems.
For three years I watched that thesis from a distance and kept recommending monoliths to clients, because that’s what worked and that’s what was accessible.
Fugu is Sakana saying: we were right. Here’s the proof.
And I think they might be right. Not entirely, not yet. But enough that I’m paying attention differently now.
The part that actually changes something for me isn’t the benchmarks. It’s the architecture.
Every workflow I’ve built assumes you pick a model and commit. You design around its strengths, you route around its weaknesses, you build guardrails for the gaps. That’s the job. That’s what I teach.
Fugu is designed for a world where you don’t commit. Where the routing layer makes the model selection decision at inference time, not at architecture time. Where if Anthropic restricts access tomorrow — or raises prices, or changes the API — the system routes around it without you rebuilding anything.
I have had people lose entire pipelines to a single API deprecation. One model goes offline, one pricing change hits, and three months of engineering work needs redoing. I have told those clients “this is the cost of building on frontier models.” I’ve accepted it as a fact of the field.
Fugu is built on the premise that this doesn’t have to be a fact.
That the right architecture makes it optional.
I don’t know if that’s fully true yet. The routing layer is opaque, which is a real problem for compliance-sensitive work.
The EU/EEA exclusion limits adoption. The latency overhead on simple tasks is genuine. If you’re making quick, single-turn API calls, orchestration adds cost with no benefit.
But for teams running long-horizon agentic workflows the kind I spend most of my time building and the case is real enough to test.
Here’s where I actually land.
I’m not changing my default recommendation tomorrow. The people I work with have specific needs, specific constraints, specific workflows that I’ve matched to specific models for real reasons.
That doesn’t go away because a new benchmark dropped.
But I’m aware, for the first time in a while, that my defaults were formed in a slightly different world. One where the frontier was clearly owned by a small number of US-based labs and the job was to choose between them wisely.
Fugu is the first production system built explicitly around the idea that that world is already ending. That frontier capability is going to keep fragmenting, keep getting restricted, keep shifting — and the teams that survive that are the ones who never bet their architecture on one provider staying accessible.
I find that argument uncomfortable to sit with, because I’ve built a lot on that bet.
I also find it hard to dismiss.
I’m testing Fugu this week on three real workflows. I’ll tell you what I actually find.
메타데이터
- post_id
- 98e0c7a62442
- slug
- when-all-recommending-claude-then-this-sakana-ai-happened-98e0c7a62442
- url
- https://medium.com/everyday-ai/when-all-recommending-claude-then-this-sakana-ai-happened-98e0c7a62442
- canonical_url
- https://medium.com/everyday-ai/when-all-recommending-claude-then-this-sakana-ai-happened-98e0c7a62442
- author_url
- https://medium.com/@singh.manpreet171900
- status
- ok
- fetched_at
- 2026-09-14 01:17:41