← Back to list

is Sakana Fugu better than Mythos and fable? The Approach is clever. The sovereignty pitch is not.

My take on the orchestration model Sakana AI shipped this week: what it actually does, where it is genuinely new, and where the marketing…

Nikhil in Neural Notions · 2026-06-22 17:40 · 0 claps · 6.7 min read paywalled
#anthropic-claude #technology #llm #artificial-intelligence #ai
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General ECO · Economy · General

is Sakana Fugu better than Mythos and fable? The Approach is clever. The sovereignty pitch is not.

My take on the orchestration model Sakana AI shipped this week: what it actually does, where it is genuinely new, and where the marketing runs ahead of the product.

is Sakana Fugu better than Mythos and fable

is Sakana Fugu better than Mythos and fable

I spend most of my working life building product on top of other people’s model APIs, so a launch like this one lands close to home. On 22 June, the Tokyo research lab Sakana AI released Sakana Fugu.

The one line version is that it is a model whose whole job is to run other models. You talk to a single endpoint, and behind it Fugu decides which models to call, splits the work between them, checks the results, and hands you back one answer.

The line that made it spread fast was bolder than the product itself: frontier capability, Sakana says, without the risk of export controls.

My short verdict is that the product is genuinely interesting and the headline is oversold. Here is the longer version.

What Sakana actually shipped

Fugu is not a framework you wire together yourself, and it is not a simple router that maps keywords to a model.

It is itself a language model, a fairly small one, trained to call a pool of larger models and even to call fresh copies of itself when a problem needs breaking into parts.

Sakana describes it as a system of many agents that behaves like a single model.

From your side you send one request to an OpenAI compatible endpoint, and the model selection, the handing off of subtasks, the checking, and the stitching back together all happen on Sakana’s side. None of that machinery shows up in your code.

There are two versions at launch.

Plain Fugu is the everyday one, tuned for speed as much as quality, the sort of thing you would drop into a coding tool or a chatbot.

Fugu Ultra coordinates a deeper set of expert models and is built for hard, many step work: research runs, reproducing papers, security assessments, patent and literature searches.

Both sit behind the same API, so moving between them does not change your integration. You can also tell Fugu to leave specific providers or models out of its pool, which matters if you have rules about where your data is allowed to go.

The part I find most credible is the research underneath it. This is not a weekend wrapper. It builds on two papers Sakana had accepted at ICLR 2026: TRINITY, an evolved coordinator that hands out Thinker, Worker and Verifier roles across a mix of models, and the Conductor, which uses reinforcement learning to work out how agents should talk to one another in plain language instead of through hand written rules.

Sakana has spent years on model merging and its AI Scientist project, so an orchestration product is a natural next step for this particular lab rather than a bandwagon move.

The benchmarks are good, and they are not a clean sweep

Here is where I want to slow down, because the numbers are being read too generously. Every figure in the launch comes from Sakana itself, and the comparison scores for rival models are those providers' own published numbers, which means the test harness and effort settings are not matched.

Read all of it as a vendor self report until someone independent runs it.

On its own headline coding test, SWE Bench Pro, Fugu Ultra scored 73.7. That puts it ahead of Opus 4.8 at 69.2, GPT 5.5 at 58.6 and Gemini 3.1 Pro at 54.2.

It also sits below Anthropic’s Fable 5 at 80.0, which happens to be the one model Sakana is loudest about matching and the one it cannot actually use.

Across the rest of the suite the picture is mixed in an honest way. Fugu Ultra leads on several reasoning and coding tests and edges Opus 4.8 on Humanity’s Last Exam by two tenths of a point.

But Fable 5 tops both SWE Bench Pro and Humanity’s Last Exam, GPT 5.5 wins the long context recall test, and Opus 4.8 takes the main cybersecurity benchmark.

On a couple of tests, plain Fugu beats Fugu Ultra, which tells you that more orchestration is not always the better choice.

The fair reading is simple. An orchestrated pool can plausibly match or beat any single model sitting inside it, and on Sakana’s own numbers Fugu belongs in the frontier conversation. Whether it matches the two models it is not allowed to include, Fable 5 and Mythos, is exactly the claim I would hold most loosely.

The sovereignty story, and why I am not buying it

This is the part that travelled.

Sakana wraps Fugu in a political argument: progress so far has come from giant single models, the future belongs to coordinated ecosystems, and leaning on one company for critical work is, in chief executive David Ha’s phrase, "a material vulnerability."

They point to the real event behind all of this. In the middle of June, national security export controls landed on Anthropic’s most capable models, Fable and Mythos, and access for organisations across a long list of countries disappeared almost overnight.

Watch full details about Fable 5:

[embed]

Because Fugu’s pool is swappable, the argument runs, if one provider goes dark the system simply routes around the gap, and that makes it a blueprint for what they call AI sovereignty.

There is a real point in here. Single vendor dependence is a genuine operational risk. Anyone who has had a model deprecated, repriced or rate limited in the middle of a build knows what a hard dependency costs, and a diverse pool you can swap is a sensible hedge.

But the word sovereignty is doing far more work than the product can support, and three problems sit right underneath it.

First, the hedge still rents its intelligence. Fugu routes around the loss of any one provider, but its capability is the pool, and the pool is other companies' models reached through their APIs. A broad restriction, rather than a single one, shrinks the pool. The resilience comes from variety, not from independence.

Second, the legal footing is unsettled. Orchestrating and reselling access to several proprietary models through one endpoint sits in a grey area of each provider’s terms of service, and that is a question every adopter inherits, not only Sakana.

Third, and this is the one I keep returning to, Fugu measures itself against the two models it is forbidden to use. Standing shoulder to shoulder with Fable 5 is a claim about a stand in, not a way to obtain Fable 5’s output.

I am not alone in this read. Elie Bakouch, a research engineer at Prime Intellect, put it bluntly on X:

a closed orchestrator sitting on top of closed models, where you used to not control the models and now do not even control which ones get used or how much, is "not AI sovereignty."

A widely shared comment on Reddit called it, until proven otherwise, a very advanced router rather than the kind of real jump in intelligence that Fable and Mythos represented.

Sakana has also not said what share of the work is carried by closed models versus open ones, and that single disclosure would settle a good deal of the argument.

Price, and the European blind spot

For individuals the pricing is approachable. Subscriptions sit at $20, $100 and $200 a month, with the higher tiers giving ten and twenty times the usage allowance, and every tier includes both models.

There is a free second month if you subscribe before the end of July. The pay as you go rate for Fugu Ultra is $5 per million input tokens and $30 per million output tokens, climbing once you push past a very large context window.

That output rate is squarely premium, in line with the pricier flagships rather than the budget tier, so the money story is really about not having to build your own orchestration rather than about cheap tokens.

The hard limit is geographic. Fugu is not available in the European Union or the wider EEA at launch while Sakana works through GDPR compliance. That is an awkward gap, because the regulated, critical infrastructure buyer the sovereignty pitch is aimed at is largely the same buyer who sits inside the European rules that currently lock them out.

What I would actually do with it

Stripped of the marketing, the important thing here is not this single product.

It is that orchestration has turned into something you can buy rather than only something you build. That now sits beside three options teams already use: aggregator style routing, the do it yourself frameworks, and the in harness workflows the model vendors ship themselves. The useful question is which of them fits the job in front of you.

If I had a messy, varied workload spanning code, reasoning and research, and I cared more about top end answer quality than about controlling every step, I would put Fugu Ultra on the shortlist and let it earn its premium, assuming my data and region rules allowed it at all. If I had a known task mix, tight cost targets and a team that can own routing and observability, I would keep the orchestration inside my own code where I can see it and tune it. And if the task is narrow and one strong model already handles it well, I would not bolt on orchestration just to feel current.

One more irony is worth sitting with. Adopting Fugu to reduce vendor dependence quietly adds a new dependence, on Sakana’s orchestrator and on whatever its pool happens to contain on a given day. That can still be a good trade. It is not an escape.

Sakana Fugu is a real, well made piece of engineering built on real research, and the direction it points in, coordinating models rather than only scaling them, is hard to argue against.

The story stapled to it, that you can buy your way out of export controls and vendor lock in, gets softer the moment you remember whose models are doing the actual work. I would benchmark it on my own traffic, wait for the independent evaluations to land, and treat the sovereignty line as marketing rather than architecture. Clever product. Oversold story.


메타데이터
post_id
b49d4b1fbe5b
slug
is-sakana-fugu-better-than-mythos-and-fable-the-approach-is-clever-the-sovereignty-pitch-is-not-b49d4b1fbe5b
url
https://pub.neuralnotions.ai/is-sakana-fugu-better-than-mythos-and-fable-the-approach-is-clever-the-sovereignty-pitch-is-not-b49d4b1fbe5b
canonical_url
https://pub.neuralnotions.ai/is-sakana-fugu-better-than-mythos-and-fable-the-approach-is-clever-the-sovereignty-pitch-is-not-b49d4b1fbe5b
author_url
https://medium.com/@nkwrites
status
ok
fetched_at
2026-06-23 06:34:20