Devin Fusion scores 61.7 on Coding Agent Index, nearly matching Claude Code at 36% less cost
ArtificialAnlys · x · 2026-09-12
Artificial Analysis independently benchmarked Devin Fusion on its release day — the first multi-model coding agent included on the Artificial Analysis Coding Agent Index (v1.5).
- Fusion CLI config: Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7, nearly tied with Claude Fable 5.1 (max with fallback) in Claude Code at 62.2.
- Cost: $7.9 per task vs $12.4 for Claude Code — 36% cheaper — with essentially flat speed (35.8 vs 34.8 min per task).
- Per-eval: 63.1 vs 64.3 on DeepSWE 1.1, 65.9 vs 64.8 on SWE-Atlas QnA, 56.1 vs 57.6 on Terminal-Bench 4.0.
- In both tested configurations, Fusion sits on the Pareto frontier of score vs cost per task.
Fusion's approach: a frontier lead model paired with a cost-efficient sidekick, retaining frontier performance while cutting costs.
More from coding & agent
- Developer streams an AI agent playing Minecraft all day via Codex, aiming to build an auto chicken farm — nickbaumann_ · 2026-09-12
- Why working with AI agents all day is exhausting: delegation adds mental load, not removes it — bendee983 · 2026-09-12
- "Codex is my main interface to the system": user says agent made Linux admin effortless — mark_k · 2026-09-12
- Alchemy's distilled project builds unified agent-friendly SDKs for 60+ SaaS services — samgoodwin89 · 2026-09-12
- RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks — echen · 2026-09-12
- Expected Parrot's universal remote cache hits 30M entries to fix AI-agent reproducibility — soumitrashukla9 · 2026-09-12