Devin Fusion hits 61.7 on Coding Agent Index, costs 36% less than Claude Code per task
ArtificialAnlys · x · 2026-09-12
Artificial Analysis published detailed per-eval results for both Devin Fusion configurations.
Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Coding Agent Index v1.5, nearly tied with Claude Code's Fable 5.1 (max, with fallback) at 62.2, while costing 36% less ($7.9 vs $12.4 per task) with essentially flat speed (35.8 vs 34.8 min/task). Per-eval: DeepSWE 1.1 63.1 vs 64.3, SWE-Atlas QnA 65.9 vs 64.8, Terminal-Bench 4.0 56.1 vs 57.6. Both configs sit on the Pareto frontier of score vs cost.
More from coding & agent
- Developer streams an AI agent playing Minecraft all day via Codex, aiming to build an auto chicken farm — nickbaumann_ · 2026-09-12
- Why working with AI agents all day is exhausting: delegation adds mental load, not removes it — bendee983 · 2026-09-12
- "Codex is my main interface to the system": user says agent made Linux admin effortless — mark_k · 2026-09-12
- Alchemy's distilled project builds unified agent-friendly SDKs for 60+ SaaS services — samgoodwin89 · 2026-09-12
- RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks — echen · 2026-09-12
- Expected Parrot's universal remote cache hits 30M entries to fix AI-agent reproducibility — soumitrashukla9 · 2026-09-12