explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Why Two Releases Instead of One
  • Fugu Max: Cheaper, Not Just Different
  • Fugu Ultra v2: The Claim That Matters
  • The Skepticism That's Owed Here
  • The Resiliency Argument Underneath the Benchmarks
  • What Comes Next
  • Related Reading on explainx.ai
← Back to blog

explainx / blog

Sakana Fugu Max and Ultra v2: Beating Opus 5 Without Calling It

Sakana AI, AI Agents, Multi-Agent Systems, Orchestration, Foundation Models

Sakana shipped Fugu Max and Ultra v2 on Sep 11, 2026. Ultra v2 claims to beat Opus 5 on Chartography without Fable 5 in its model pool.

Sep 15, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Sakana Fugu Max and Ultra v2: Beating Opus 5 Without Calling It

Sakana AI didn't wait for a bigger single model. On September 11, 2026, at 7:45 AM, @SakanaAILabs announced two new releases in its Fugu orchestration line — Fugu Max and Fugu Ultra v2 — pushing the same idea from its original Fugu launch in two directions at once: cheaper "good enough" routing, and higher peak capability, without merging the two goals into one tradeoff.

Sakana's own summary line for the pair is worth quoting directly, because it states the thesis cleanly: "Fugu Max expands the Pareto frontier outward. Fugu Ultra v2 pushes it upward." Read together with the most interesting claim in the announcement — that Ultra v2 beats Opus 5 and Fable 5 on two benchmarks without including those models in its own agent pool — this release is less a spec bump than a bet on orchestration itself as the unit of progress, not any single model inside it.

TL;DR

table · 2 cols
QuestionAnswer
What launched?Fugu Max (cost-efficiency tier) and Fugu Ultra v2 (peak-capability tier), announced Sep 11, 2026
What's the core idea?Route each task to the leanest capable model in a pool, instead of betting on one frontier model
What's in the pool?Sakana's largest pool yet of open-weight and specialized models, including the NVIDIA Nemotron family
Does it call Opus 5 or Fable 5?Sakana says no — Fugu Ultra v2 excludes Fable 5, Fable 5.1, and GPT-6-Astra from its own pool
Headline benchmark claimChartography: 48.3 (Fugu Ultra v2) vs. 27.3 (Opus 5), per Sakana's own reporting
Cost claim for Fugu Max"Within striking distance" of frontier models at 2-6x lower cost
Where do you try it?sakana.ai/fugu
Independently verified?No — these are Sakana's self-reported numbers as of publication

Why Two Releases Instead of One

Most model vendors ship a single upgrade and let users decide where it sits on the cost-vs-capability curve. Sakana's framing rejects that as the wrong mental model entirely. In Sakana's telling, the industry still treats model choice as a static menu — you pick GPT-6-Astra or Opus 5 or a cheaper open model, and you live with wherever that lands you on cost and capability. Fugu's premise, going back to its original June 2026 launch, is that this menu should be dynamic: a system sitting in front of the model pool that decides, task by task, which model is the leanest one capable of doing the job.

Fugu Max and Fugu Ultra v2 split that system's improvement into two separate axes rather than one blended tradeoff:

  • Fugu Max expands the frontier outward — more capability per dollar, by orchestrating a larger pool of open-weight and specialized models, including the NVIDIA Nemotron family, and routing dynamically to whichever is cheapest for a given task.
  • Fugu Ultra v2 pushes the frontier upward — higher peak capability, aimed at benchmark parity or superiority against frontier closed models, independent of cost.

This is the same "Pareto frontier" language that shows up across cost/capability comparisons of frontier models — the curve of models where you can't get more capability without paying more, and can't pay less without losing capability. Sakana's bet is that an orchestration layer, not a single model, is what actually moves that curve, because it can mix and match rather than being stuck wherever one training run landed.

Fugu Max: Cheaper, Not Just Different

Sakana claims Fugu Max delivers performance "within striking distance" of elite, frontier-class models at 2-6x lower cost. The mechanism is the same dynamic routing Fugu has used since launch — internally deciding whether a request needs a heavyweight specialized model or can be handled by a leaner one — but applied to what Sakana describes as its largest model pool yet.

The explicit inclusion of the NVIDIA Nemotron family is notable. Nemotron models are open-weight and were purpose-built by NVIDIA for efficient inference at scale, which makes them a natural fit for an orchestrator whose entire value proposition is routing cheap tasks to cheap models. It's the same logic behind GitHub Copilot's HydraFusion, which made a similar bet that model selection — not model size — is where near-term cost savings live inside coding assistants.

The 2-6x figure is Sakana's own claim, not an independently reproduced benchmark. Cost multipliers like this depend heavily on task mix — a workload that's mostly simple lookups will route to cheap models most of the time and post a large multiplier; a workload full of genuinely hard reasoning tasks will route to expensive models more often and post a much smaller one. Treat "2-6x" as a range that depends entirely on what you're actually asking it to do.

Fugu Ultra v2: The Claim That Matters

Fugu Ultra v2 is where the announcement gets substantive rather than promotional. Sakana states that Ultra v2 outperforms Claude Opus 5 and Fable 5 on a benchmark it calls Chartography, and outperforms models costing 3-5x more per token on a benchmark called DeepSWE.

A follow-up post from Sakana added specific numbers:

  • Fugu Ultra v2 ranks #1 on 5 of 8 "hard benchmarks" Sakana tested it against, including DeepSWE, Chartography, and one called Toolathon.
  • DeepSWE score: 74.3.
  • Chartography score: 48.3, versus Opus 5's reported 27.3 on the same benchmark.

Here's the part worth reading twice. Sakana explicitly states that Fugu Ultra v2 achieves these scores without including Fable 5, Fable 5.1, or GPT-6-Astra in its own orchestrated agent pool. That distinction is Sakana pre-empting the single most obvious objection to any orchestrator claiming to "beat" a frontier model: that it isn't actually beating it at all, it's just quietly calling the expensive model under the hood and taking credit for its answer whenever the task gets hard.

If that objection held, Fugu Ultra v2's benchmark win would be meaningless — a routing layer that calls Opus 5 whenever Opus 5 would do better isn't a competing system, it's a wrapper. By stating the exclusion explicitly, Sakana is claiming the win comes from routing across open and specialized models only — a genuinely different claim than "our system knows when to phone a friend." Whether that holds up under independent scrutiny is a separate question from whether Sakana is being clear about what it's claiming, and on the second point, this framing is at least unambiguous.

DeepSWE and Toolathon both lean toward agentic, tool-use-heavy tasks — the kind that decompose naturally into subtasks a router can split across specialized models. That's consistent with where orchestration systems have historically shown their strongest gains: not raw single-shot reasoning, but multi-step work where different steps genuinely benefit from different models.

The Skepticism That's Owed Here

None of the numbers above have been independently verified as of this writing. Chartography, DeepSWE, and Toolathon are benchmarks named in Sakana's own release posts, run and scored by Sakana. That's standard practice for any lab announcing new numbers — every model vendor's launch-day benchmark table is self-reported until someone outside the company reruns it — but it means the 48.3-vs-27.3 Chartography gap and the "#1 on 5 of 8" claim are Sakana's own scoring, not a neutral third party's.

This isn't a new pattern for Fugu specifically. When the original Fugu launched in June 2026, Sakana's published benchmarks looked strong against Fable 5 and Mythos — and within 24 hours, independent testers led by Ethan Mollick reported real-world creative-coding runs that took 30 minutes and produced output that didn't match Fable 5 in practice, despite the benchmark parity on paper. Sakana later added Fugu-Cyber, which posted its own strong cyber-defense numbers, and a Sakana Chat surface bundling Fugu with a new Namazu model and in-chat code execution — each shipped with its own confident benchmark table.

That history doesn't mean Fugu Max and Ultra v2's numbers are wrong. It means the sensible default, same as with any lab's self-reported launch numbers, is to wait for independent replication — or better, to run your own task against it — before treating a benchmark slide as a production decision.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The Resiliency Argument Underneath the Benchmarks

Buried under the benchmark claims is Sakana's actual architectural philosophy, and it's the part worth taking seriously regardless of how the numbers hold up. Because Fugu orchestrates a swappable pool of open and specialized models rather than depending on any single frontier model, Sakana argues it protects users from vendor lock-in, API revocations, and sudden service cutoffs.

This isn't hypothetical for Sakana's audience. Anthropic's Fable 5 and Mythos models had export controls imposed on them earlier in 2026, and organizations that built critical infrastructure on those APIs found access could shift or disappear overnight due to regulatory decisions outside their control. An orchestration layer that can swap out any single model in its pool — because no individual model is load-bearing for the whole system — is a genuinely different resiliency posture than betting a production stack on one vendor's API staying reachable.

This is the strongest version of Sakana's pitch, and it's architecturally distinct from how most teams currently hedge against vendor risk, which is usually a manually maintained fallback path to a second provider. Fugu's bet is that the fallback decision should be automated and continuous, not a break-glass procedure — closer in spirit to how Perplexity's own GLM orchestrator work treats model choice as a live cost-and-availability decision rather than a fixed configuration.

What Comes Next

Sakana shipped two other Sakana AI research announcements the same week — including a separate, unrelated PC-ALM predictive-coding paper on training deep networks without backpropagation — which is a reminder that Sakana is running multiple research tracks in parallel, not just iterating on one product line. Fugu Max and Ultra v2 are squarely the product side: something a team evaluating model spend or vendor risk this week can actually try, at sakana.ai/fugu, rather than a research result to file away.

The practical test is the same one that applied to the original Fugu launch: benchmarks rank systems, but they don't tell you whether your specific workload — your prompts, your tool calls, your latency budget — routes the way Sakana's test suite did. Run it on something that matters before deciding whether the Pareto frontier actually moved for you.

Related Reading on explainx.ai

  • Sakana Fugu: One Model API to Orchestrate All the Others — the original June 2026 launch and Mollick's Harbor bench findings
  • Fugu-Cyber: Sakana AI Orchestrates Frontier Models for Cyber Defense — the cyber-defense-specific Fugu endpoint
  • Sakana Chat: Fugu, New Namazu, and In-Chat Code Execution — Fugu's no-API consumer surface
  • Sakana Fugu in Claude Code: Setup, Pricing, and Limits — using Fugu inside a coding harness
  • GitHub Copilot HydraFusion: model selection vs. model orchestration — a mainstream coding assistant shipping the same routing idea
  • Perplexity Computer's GLM 5.2 orchestrator — another live cost-vs-capability routing decision in production
  • NVIDIA Nemotron 3 Ultra: 550B MoE open-weight model — inside the model family powering part of Fugu's pool
  • Sakana AI's PC-ALM: training deep nets without backpropagation — Sakana's other September 2026 research release
  • Browse the full LLM directory — compare Fugu Max, Fugu Ultra v2, Opus 5, and every other tracked model
  • Explore AI agents — autonomous systems that pair naturally with orchestration APIs like Fugu

This post reflects Sakana AI's own September 11, 2026 announcement of Fugu Max and Fugu Ultra v2. Benchmark figures (Chartography, DeepSWE, Toolathon) are Sakana's self-reported numbers as of publication and have not been independently verified — treat them as a starting hypothesis, not a settled result, and test against your own workload before migrating anything production-critical.

Spotted something out of date? Let us know.

People in this article

  • Ethan Mollick →Associate professor of management at Wharton
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Jun 22, 2026

Sakana Fugu: One Model API to Orchestrate All the Others

Sakana AI's Fugu Ultra launched June 22 with bold benchmark claims against Fable 5 and Mythos. Within 24 hours, Ethan Mollick and other testers reported 30-minute shader runs, ~$6 per demo, and output that does not match Fable in real use — despite strong published scores. Here is what the Harbor bench reveals.

Jul 21, 2026

Fugu-Cyber: Sakana AI Orchestrates Frontier Models for Cyber Defense

Sakana AI launched Fugu-Cyber on July 21, 2026, extending its multi-model orchestrator into vulnerability verification and threat-intelligence detection. The benchmark scores are strong, but Sakana's more important argument is that enterprises need verification harnesses and human expertise, not merely access to a frontier cyber model.

Sep 5, 2026

GitHub Copilot HydraFusion: Model Orchestration Over Model Selection

Satya Nadella tweeted about Project HydraFusion on September 4, 2026 — a GitHub Copilot research preview that routes coding tasks across drafting, critique, and escalation models instead of running one model end to end. Here's what the official post actually says, how the orchestration works, and why "model orchestration" is becoming the next competitive axis for agent harnesses.