Coding Agent Index: Sonnet 5.5 tops at 68 but costs $14.19/task vs GPT-6.1 Sol's $1.04
ArtificialAnlys · x · 2026-10-02
Artificial Analysis updated its Coding Agent Index (DeepSWE v1.1, Terminal-Bench 4.0, SWE-Atlas-QnA) with three new frontier models near the top:
- Claude Sonnet 5.5 (max) in Claude Code takes first place at 68, but with the highest measured cost at $14.19 per task
- Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84/task — less than half of Sonnet's cost; note it uses Google's promotional pricing and Argon isn't publicly available yet
- GPT-6.1 Sol (xhigh) in Codex scores 63 at just $1.04/task, roughly one-sixth of Argon's cost
Takeaway: performance is tightly clustered while costs span an order of magnitude, making cost-per-task the key differentiator.
More from coding & agent
- Snorkel AI lands five NeurIPS papers; hardest long-horizon agent tasks see <1% pass rate — ajratner · 2026-10-02
- cronstable: an open-source job scheduler with a built-in MCP server for agent-driven triage — ptweezy · 2026-10-02
- What happens when four AI agents update the same file? Testing Git worktrees vs AgentWS — pilver7 · 2026-10-02
- Claude Given 12 Hours Autonomously Produced The Clodyssey — FinanceYF5 · 2026-10-02
- Hebbrix Launches Hosted MCP Memory Endpoint, Warns of Write-to-Searchable Gap — Separate_Sand8265 · 2026-10-02
- Agent hits first roadblock: safety filter forces manual send, says Greg Mushen — gregmushen · 2026-10-02