Devin Fusion benchmarks: frontier lead model plus cheap sidekick cuts coding agent costs
ArtificialAnlys · x · 2026-09-12
Artificial Analysis independently benchmarked Devin Fusion, Cognition's newly released multi-model coding agent and the first of its kind on their Coding Agent Index.
Fusion pairs a frontier lead model (Claude Fable 5.1 or GPT-6 Astra, both xhigh) with a cost-efficient SWE-2 (medium) sidekick. The Fable config scores 62 on the Index v1.5, while the Astra config scores 59 but is 43% cheaper and 31% faster — retaining frontier-level performance at lower cost.
More from coding & agent
- Developer streams an AI agent playing Minecraft all day via Codex, aiming to build an auto chicken farm — nickbaumann_ · 2026-09-12
- Why working with AI agents all day is exhausting: delegation adds mental load, not removes it — bendee983 · 2026-09-12
- "Codex is my main interface to the system": user says agent made Linux admin effortless — mark_k · 2026-09-12
- Alchemy's distilled project builds unified agent-friendly SDKs for 60+ SaaS services — samgoodwin89 · 2026-09-12
- RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks — echen · 2026-09-12
- Expected Parrot's universal remote cache hits 30M entries to fix AI-agent reproducibility — soumitrashukla9 · 2026-09-12