2026-09-10

Sakana's Fugu Max and Ultra v2 Orchestrate Without a Single Frontier Model — and Score Opus 5 Nothing Like DeepSeek Did Yesterday

AIModelsInfrastructure🌍 Asia

Sakana AI announced two new tiers of its Fugu orchestration system today: Fugu Max, aimed at cost efficiency, and Fugu Ultra v2, aimed at peak capability. Both build on the idea this blog covered when Sakana shipped Fugu-Cyber in July: Fugu isn't a model, it's a system that routes tasks across a pool of other models and behaves like a single frontier model from the outside. Sakana's own release post lays out the lineage in full: an April beta, general availability plus "Fugu Ultra v1" in June, Fugu-Cyber in July, a consumer chat product and the start of Sakana's NVIDIA Nemotron integration in August, and Max plus Ultra v2 today — so the Nemotron partnership behind today's pool is a month old, not brand new. What's new this time is the pool itself — Sakana describes Fugu Max as orchestrating "our largest pool of open and specialized models to date," dynamically routing each task to "the leanest model capable of solving them" rather than defaulting to the biggest one available. A companion thread post sharpens that framing further: an "unprecedented mix of open and closed models" — the pool isn't open-weights-only, it specifically excludes a short, named list of current frontier flagships rather than closed models generally.

The split, and what it actually costs

Fugu Max is priced at $2 per million input tokens and $6 per million output tokens; Fugu Ultra v2 is priced at $5 input and $30 output — a 2.5x gap on input and a 5x gap on output between the two tiers, per pricing trackers covering the release, with Sakana's own post confirming the $2/$6 Max figures directly. Sakana's own framing for Max is "performance within striking distance of elite models at two to six times lower cost," measured, by its own account, "among frontier models in a similar price range" — a price-banded comparison, not a claim of beating every model regardless of tier. By the same post, Max's output pricing runs 40-60% below Claude Sonnet 5, GPT 5.6 Terra, and Kimi K3, and Sakana separately claims Max achieves the best score among similarly-priced models on six of ten tracked benchmarks (Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish) and expands the cost-performance frontier on seven of the ten. Sakana did publish an eight-benchmark grid for Ultra v2 (see below), but it shows raw capability scores, not a cost-normalized comparison — so the specific "two to six times" multiplier for Max still isn't a number this post can trace to a shown calculation.

Beating Opus 5 and Fable 5 without either of them in the pool

The more interesting claim is architectural, not just economic: Sakana states plainly that Fugu Ultra v2 does all of this "without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool" — meaning its headline results come from orchestrating older, open, and specialized models rather than calling out to the current frontier and repackaging the answer. A separate post in the thread makes the rationale explicit: "A single frontier model can be restricted. A single API can be revoked. But a swappable, orchestrated pool routes around disruption by design. Model resiliency is not a backup feature. It is the architecture." That's a real differentiator from a router that just forwards each query to whichever frontier model handles it best (roughly what Sakana's original, simpler Fugu did) — and Sakana's own diagram illustrates the resulting pitch abstractly, plotting Fugu Max and Fugu Ultra above a "frontier formed by single models" on an unlabeled performance-vs-cost axis. It's a schematic, not a benchmarked chart, but it's the clearest visual statement of the thesis in the launch materials. Per Sakana's own benchmark suite, Ultra v2 is best or joint-best on five of eight tracked benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon — and in the top two on seven of eight, benchmarked directly against Opus 5, Fable 5, Fable 5.1, GPT-5.6 Sol, Kimi K3, and GPT-6-Astra even though the latter three are excluded from its own pool. Its DeepSWE score, 74.3, lands almost exactly where this blog's coverage of DeepSeek's V4.1 Flash launch put Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0) on the same-named benchmark one day earlier — a genuinely reassuring cross-check, since two labs' independently reported numbers for the same task cluster within a point of each other.

The Chartography number that doesn't cross-check at all

One benchmark doesn't cluster. Sakana's own published chart puts Fugu Ultra v2 at 48.3 on Chartography, against 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5. This blog's post on DeepSeek's official V4.1 Flash launch, published the day before, reported Claude Opus 5 scoring 84.0 on a benchmark shown on DeepSeek's own launch chart under the name "Chartography (w/tools)." That's not a small discrepancy — it's the difference between Opus 5 finishing near the bottom of an eight-way comparison and finishing near the top of a different one, on what reads as the same named evaluation. The likeliest explanation is that "Chartography" isn't a standardized, shared benchmark the way DeepSWE apparently is: different labs may be running different question sets, different tool-access configurations, or different scoring rubrics under the same name, and DeepSeek's "(w/tools)" qualifier suggests exactly that kind of configuration difference — Sakana's own chart doesn't note whether its Chartography run used tools at all. Nothing here proves either lab fabricated a number — but it's a clean demonstration of why a shared benchmark name, on its own, doesn't make two labs' self-reported charts comparable, and why a single-source benchmark claim is worth treating as provisional until an independent evaluator runs the same models under the same conditions.

A different kind of "sovereignty" claim

Sakana closes its own release post with a line that reframes the whole launch: Fugu provides "the resilient, vendor-agnostic infrastructure required for true AI sovereignty." That's a notably different definition of sovereignty than the one this blog has tracked all month in Mistral's €3 billion Series D or France's economy minister warning that European AI can't rest on one national champion — those are arguments about which country's capital and which company controls a model. Sakana's version is about vendor independence: a pool of swappable models that no single provider can restrict, revoke, or cut off "due to vendor lock-in, API revocations, geopolitical turbulence, and sudden service cutoffs," regardless of whose flag is on the model weights. Both are real problems worth solving, but calling vendor-agnostic orchestration "true" sovereignty, in the same month as several stories about which government and which conglomerate actually owns the underlying compute, is a claim doing quiet work of its own.

What to expect next

  • Watch for independent evaluation of Fugu Ultra v2 on Chartography specifically, given how far its reported Opus 5 score diverges from this blog's own reporting on the same-named benchmark just one day earlier.
  • Watch whether Sakana publishes the full per-benchmark table behind "two to six times lower cost." The compressed range is a marketing summary; the underlying comparison, if released, is the actual test of the claim.
  • Watch how "no frontier models in the pool" holds up as a durable pitch, given NVIDIA's own Nemotron models are explicitly built to serve as cheap workers inside exactly this kind of orchestration system — Sakana isn't just avoiding frontier models, it's leaning on hardware-vendor-subsidized alternatives built for this specific role.
  • Watch whether Fugu-Cyber, Sakana's security-specific variant, gets a similar Max/Ultra split, or whether this generational jump stays confined to the general-purpose orchestrator.