Sakana's Fugu Ultra v2 tops 5 of 8 benchmarks by orchestrating open models, no frontier models
SakanaAILabs · x · 2026-09-12
Alongside the Fugu Ultra v2 launch, Sakana AI published a blog laying out the cost-performance two-axis strategy behind its orchestration systems.
- Benchmarks: best or joint-best on 5/8 benchmarks (DeepSWE 74.3, Chartography 48.3 beating Opus 5 and Fable 5, Toolathon, GDP.pdf, SWEFish) — achieved without Fable 5, Fable 5.1, or GPT-6-Astra in the agent pool
- One architecture, two missions: Fugu Max maximizes output quality at minimum cost, expanding the Pareto frontier; Fugu Ultra v2 pushes peak capability on complex multi-step tasks
- Core thesis: the future belongs to systems that know which machinery to deploy per task, not those that indiscriminately call the largest model; an orchestration layer can rival single frontier models without depending on proprietary ones
- Trajectory: April beta proved multi-agent orchestration works as a unified foundation model; June GA and Ultra v1 showed the layer could match closed frontier models
Note: orchestration tokens are billed as standard input/output, so real cost structure requires careful accounting.
More from coding & agent
- OpenAI details GPT-6 Astra quality fixes: misfiring legacy skills, experiment hit 4-5k users — rohanpaul_ai · 2026-09-13
- A surprising Codex prompt trick inspired by 'Staying Human in the Age of AI' — daniel_mac8 · 2026-09-13
- Cooking With Blender MCP and Astra: AI Driving 3D Workflows Directly — sidahuj · 2026-09-13
- Bind agent approvals to proposal hashes: any change should invalidate them — arthaudm · 2026-09-13
- 10 GitHub repos that give AI agents superpowers: browsers, memory, GUI control — Shruti_0810 · 2026-09-13
- User: GLM-5.3 coding subscription beats Grok, but pair it with Opencode — thatroblennon · 2026-09-13