Sakana Fugu: multi-agent orchestration as a model
- What happened
- Sakana AI launched Fugu, a multi-agent orchestration model that coordinates frontier models behind a single API — two tiers, OpenAI-compatible, backed by two ICLR 2026 papers.
- Why it matters
- Fugu Ultra beats GPT 5.5, Opus 4.8, and Gemini 3.1 Pro on SWE Bench Pro and Humanity's Last Exam at $5/$30 per 1M tokens — it's a new category of model that competes on coordination, not parameter count.
- What to do
- Try Fugu in any OpenAI-compatible client as a complement to your stack; don't switch your primary model yet — EU teams should wait for GDPR compliance.
Sakana AI shipped Fugu on June 22, 2026 — a multi-agent orchestration model that coordinates frontier models behind a single OpenAI-compatible API. Fugu Ultra posts scores that beat GPT 5.5, Opus 4.8, and Gemini 3.1 Pro on SWE Bench Pro (73.7 vs. 58.6) and Humanity's Last Exam (50.0 vs. 41.4), at $5 per 1M input tokens and $30 per 1M output tokens. Our verdict: Watch. Two ICLR 2026 papers back the architecture, but it's day one — no production track record, no EU availability, and the pricing introduces a new mental model for cost calculation. We'll test and return with a recommendation.
What happened
Sakana AI launched Fugu at sakana.ai/fugu(opens in new tab), grounded in two ICLR 2026 papers: TRINITY (an evolved LLM coordinator) and Conductor (reinforcement-learning-based agent orchestration) (Sakana AI, 2026). Instead of prescribing fixed team roles or hand-designed workflows, Fugu learns to dynamically assemble agents from a pool, assign them Thinker, Worker, or Verifier roles, and coordinate them through collaboration patterns that humans wouldn't design.
Two tiers ship:
- Fugu — balanced latency and quality. Designed for everyday coding and interactive work. You can configure which models enter the pool, opting out specific providers to meet privacy and compliance constraints.
- Fugu Ultra — deeper agent pool, longer response times, optimized for complex multi-step reasoning. Fixed, proprietary pool — you can't see which models it used. Target use cases: paper reproduction, Kaggle competitions, cybersecurity analysis, patent landscape research.
Both tiers expose a single OpenAI-compatible endpoint. Point any existing client at the Fugu API with your key — no SDK migration required.
Benchmarks that demand attention
Fugu Ultra's scores on rigorous engineering and reasoning benchmarks beat or tie publicly available frontier models:
| Benchmark | Fugu Ultra | Opus 4.8 | Gemini 3.1 Pro | GPT 5.5 |
|---|---|---|---|---|
| SWE Bench Pro | 73.7 | 69.2 | 54.2 | 58.6 |
| LiveCodeBench | 93.2 | 87.8 | 88.5 | 85.3 |
| Humanity's Last Exam | 50.0 | 49.8 | 44.4 | 41.4 |
| GPQA-D | 95.5 | 92.0 | 94.3 | 93.6 |
The qualitative demos are harder to dismiss: Fugu solving blindfold chess against a 2100-Elo Stockfish engine, reproducing a research paper end-to-end over roughly 4 hours, driving a full security assessment from a single scoped instruction, and generating a working mechanical iris in CAD where frontier baselines produced broken linkages or crashed entirely (Sakana AI, 2026).
Why it matters
Fugu isn't another LLM — it's a new category. It doesn't compete on parameter count; it competes on coordination. The model learns which agents to deploy for which subtask and how to interpret their outputs, producing results that exceed any single model in the pool.
The pricing model is equally novel. Fugu Ultra charges a fixed $5/$30 per 1M tokens (input/output), with higher rates above 272K context ($10/$45). Cached input runs $0.50/$1.00. Fugu's pricing is a single blended rate based on the most expensive model in your configured pool — no stacking of per-model fees, no surprises. Subscription plans start at $20/month (Standard), scaling to $100/month (Pro, 10× usage) and $200/month (Max, 20× usage). Early adopters who subscribe by end of July 2026 get a free second month.
Two limits matter now. First, Fugu is not available in EU/EEA member states while Sakana works toward GDPR compliance. Second, the Fugu Ultra pool is fixed and proprietary — you can't audit which models handled your query, which matters for enterprise compliance and debugging.
What changes for you
Fugu is a genuine complement to your current model stack, not a replacement. The architecture is novel enough to deserve a test drive — but it hasn't earned trust yet.
- Try Fugu (not Ultra) first. The balanced tier gives you the orchestration feel without Ultra's higher latency and cost. Drop it into any OpenAI-compatible client.
- Keep your primary model. The benchmarks are impressive, but this is a v1 product with zero production history. Use Fugu alongside your current setup.
- EU/EEA teams: wait. Sakana explicitly blocks service in those regions pending GDPR compliance.
- Watch for our test results. We'll run Fugu through our evaluation pipeline and return with a recommendation.
FAQ
What makes Fugu different from using multiple models yourself? Fugu learns coordination strategies through training, not manual prompt engineering. TRINITY and Conductor show it discovers non-obvious agent combinations and role assignments that outperform hand-designed workflows — you get the benefit without the engineering overhead (Sakana AI, 2026).
Can I see which underlying models Fugu Ultra used? No. The model selection and routing is proprietary. If auditability matters for your use case, stick with the base Fugu tier, where you control the pool from the console settings.
How does Fugu's pricing compare to using frontier models directly? Fugu Ultra's $5/$30 compares favorably to direct frontier API pricing, especially considering you're getting multi-model orchestration behind a single endpoint. For Fugu, you pay one blended rate — never a sum of per-model fees — based on the most expensive model in your pool.
What to do
- 1 Try Fugu in any OpenAI-compatible client — same API, no SDK migration
- 2 Keep your primary model; use Fugu as a complement until it proves itself in production
- 3 Wait for EU/EEA availability if you operate in those regions
- 4 Watch for our test results — we'll return with a recommendation
Affected tools & models
Never need to catch up again
The weekly delta — only verdict changes and act-now items. No digest filler.