← Back to list

The AI That Commands Other AI Just Matched Claude’s Best Model

Conductor architecture makes billion-dollar training runs look optional.

Muhamed Fazal PS in Ai-Ai-OH · 2026-06-24 09:16 · 0 claps · 3.9 min read paywalled
#claude #fable #ai #artificial-intelligence #chatgpt
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General 🏛️ · Architecture 🥊 · Combat Sports

The AI That Commands Other AI Just Matched Claude’s Best Model

Conductor architecture makes billion-dollar training runs look optional.

Generated by Chatgpt

Generated by Chatgpt

You believe the only path to frontier AI is spending billions on bigger models. Every major lab, from Anthropic to OpenAI to Google, is racing to train the single largest neural network possible. The assumption is simple: bigger model equals better performance.

Sakana AI just proved that assumption wrong. Their new system, Fugu, matches Claude Fable 5 and beats GPT-5.5 on hard benchmarks by doing something none of the giants do: it commands other models instead of competing with them.

As I explored in my analysis of the AI systems agencies sell for $1k-$4k, the real cost of AI has always been in the orchestration, not the model itself. One of Fugu’s co-founders wrote the original Transformer paper, the blueprint every AI model today runs on. He could have joined any frontier lab. Instead, he chose to build an orchestra conductor for AI.

If you are not a premium member then you can read this article for free here:

What is Fugu actually doing?

Think of it this way. You send one request to one API. Behind the scenes, Fugu decides: should I answer this myself, or should I assemble a squad of specialist models? It picks the right models, delegates tasks, verifies results, and merges everything into one answer. You never see the complexity.

Sakana launched two versions. Fugu handles fast, everyday coding and chat. Fugu Ultra tackles hard, multi-step problems like AI research, paper reproduction, cybersecurity analysis, and patent search. Both use the same conductor architecture.

The key insight is that Fugu does not try to be the best at everything. It tries to be the best at picking which model is best at each specific task. That is a fundamentally different engineering problem, and it is one that scales differently than training larger models.

This reminds me of how I approached editing videos with AI using Claude Code and Remotion — the orchestration layer mattered more than the underlying model.

The biggest AI companies are racing to build one giant model. Fugu flips the script: an LLM that commands a pool of the world’s best models, choosing who does what.

The benchmarks that matter

Sakana’s claims are specific. Fugu Ultra matches Anthropic’s Fable 5 and Mythos Preview on the hardest engineering, science, and reasoning benchmarks. It beats Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5 on tasks like AutoResearch, mechanical design, and financial forecasting.

These are not cherry-picked numbers. AutoResearch requires an AI to read papers, form hypotheses, design experiments, and execute them. Mechanical design requires spatial reasoning and engineering judgment.

Financial forecasting requires pattern recognition across noisy data. Fugu Ultra handles all three by routing each subtask to the model best suited for it.

The benchmark results suggest that model orchestration can match or exceed monolithic training on complex, multi-step tasks. This does not mean Fugu is better at everything. It means the conductor approach is competitive at the frontier.

Why this changes everything

Three reasons this matters more than another billion-dollar training run.

First, export controls become irrelevant. When a provider restricts access to a frontier model, Fugu’s agent pool is swappable. It routes around the restriction by using alternative models. You cannot sanction a conductor.

Second, cost drops dramatically. Training a frontier model costs hundreds of millions. Fugu uses existing models as building blocks. The infrastructure investment is in orchestration logic, not GPU clusters. Sakana’s approach could democratize access to frontier-level performance for organizations that cannot afford to train their own models.

Third, it gets better automatically. Sakana says Fugu will naturally grow by incorporating newer, more efficient models, including their own. Every time a new model launches, Fugu can add it to the pool without retraining anything. This creates a compounding advantage that monolithic models cannot match.

The expert perspective

The Transformer paper authors understand model architecture better than anyone alive. When one of them chooses to build a conductor system instead of another frontier model, that tells you something about where the field is heading.

Industry experts have noted that the “model of models” approach mirrors how software engineering evolved. Nobody builds every library from scratch anymore. They compose existing components. Fugu applies the same logic to AI models.

The evidence hierarchy here is strong. We have benchmark data from Sakana, independent verification from beta testers, and the credibility of a team that includes Transformer paper authors. This is not an unverified claim from an unknown startup.

What this means for you

If you are building AI applications, Fugu’s approach suggests a shift. Instead of picking one model and optimizing around its weaknesses, you might soon compose multiple specialists. The API stays simple. The complexity moves to the orchestration layer.

For developers, this could mean better performance at lower cost. For enterprises, it could mean vendor lock-in becomes less dangerous. For the AI industry, it challenges the assumption that only the biggest labs can compete at the frontier. As I demonstrated in my guide on building a self-optimizing video engine with Claude, the real power comes from composing multiple AI capabilities, not from using a single model for everything.

The question is not whether Fugu’s approach works. The benchmarks say it does. The question is whether the giants will adapt, or keep building bigger models while the conductor architecture quietly takes over. You can explore Sakana’s research directly on their ArXiv publications or check their official GitHub repository for implementation details.

Environment setup

API Access: Sakana Fugu API (subscription + pay-as-you-go) Models: Fugu (fast), Fugu Ultra (max quality) Compatible with: OpenAI SDK format Pricing: Subscription for everyday use, pay-as-you-go for enterprise

Support this work: Buy Me a Coffee

Connect: MediumGitHub


메타데이터
post_id
a112eab5c874
slug
the-ai-that-commands-other-ai-just-matched-claudes-best-model-a112eab5c874
url
https://medium.com/ai-ai-oh/the-ai-that-commands-other-ai-just-matched-claudes-best-model-a112eab5c874
canonical_url
https://medium.com/ai-ai-oh/the-ai-that-commands-other-ai-just-matched-claudes-best-model-a112eab5c874
author_url
https://medium.com/@muhamedfazalps7
status
ok
fetched_at
2026-06-25 07:00:49