GPT-6 Sol and Claude Opus 5.5 tie at 97% on Next.js agent evals, with 3x cost gap
pvncher · x · 2026-09-23
Vercel's updated Next.js Agent Evals (Sept 22, 2026) measure how well AI coding agents complete real Next.js tasks.
Tied at the top: GPT-6 Sol (Codex) and Claude Opus 5.5 (Claude Code) both hit 97%, matching Claude Fable 5.1. Among the three, Opus 5.5 has the lowest average cost ($0.234/task) while Fable 5.1 is the priciest ($0.722).
Other highlights:
- Grok 4.7 (OpenCode) hits 94% at just $0.109 — standout value
- Kimi K3, GLM 5.2 and Claude Sonnet 5 sit at 81%, but all reach 97% once AGENTS.md is added, showing how much project context files matter
- MiniMax M3 is cheapest at $0.053/task; Cursor Composer 2.5 close behind ($0.062)
- Historical data shows rapid progress: Claude Opus 5 scored 94% at $1.96 in September vs. Opus 4.6's 58% at $0.204 in April
The poster argues that with per-task costs now under $0.25, piecemeal evals are saturated — agents now run for hours.
More from coding & agent
- Distill Lock: context as a build artifact with byte-identical, drift-proof, offline-verifiable output — _akpiper · 2026-09-23
- Open-source Agentic Engineering Guide: 10 parts, 33 chapters on shipping AI agents to production — _akpiper · 2026-09-23
- Learning to code still matters: Motivation + Knowledge + Skills + AI beats Motivation + AI — iamKierraD · 2026-09-23
- Low-VRAM user claims applying a Qwen-style chat template fixes Ternary Bonsai 2's tool calling — needthosepylons · 2026-09-23
- pi-voice: Handy creator ships a voice extension for Earendil Pi — mitsuhiko · 2026-09-23
- Practical guide: using Gemini API skills with your coding agents — patloeber · 2026-09-23