GPT-6 Sol and Claude Opus 5.5 tie at 97% on Next.js agent evals, with 3x cost gap

pvncher · x · 2026-09-23

Vercel's updated Next.js Agent Evals (Sept 22, 2026) measure how well AI coding agents complete real Next.js tasks.

Tied at the top: GPT-6 Sol (Codex) and Claude Opus 5.5 (Claude Code) both hit 97%, matching Claude Fable 5.1. Among the three, Opus 5.5 has the lowest average cost ($0.234/task) while Fable 5.1 is the priciest ($0.722).

Other highlights:

The poster argues that with per-task costs now under $0.25, piecemeal evals are saturated — agents now run for hours.

Original post →

More from coding & agent

coding & agent channel →