Vercel Releases Next.js AI Agent Eval: Kimi K3 and Claude Tie at the Top
evilrabbit_ · x · 2026-08-01
Vercel has released its AI Agent Evaluations for Next.js, testing major models on code generation and migration tasks based on success rate, average duration, and cost.
- Top Tier: Kimi K3, Claude Fable 5 (high), and Cursor Composer 2.5 all achieved a 92% base success rate. Kimi K3 is highly cost-effective ($0.141), while Cursor Composer 2.5 is the fastest and cheapest ($0.046).
- The AGENTS.md Boost: The benchmark highlights the impact of AGENTS.md, which bundles Next.js documentation for agents. With this context, most models (e.g., GLM 5.2, Grok 4.5) saw their success rates jump to 96%.
- Cost vs. Performance: GPT 5.5 Pro took 771s and cost $18.21 but only achieved an 83% success rate. Open-source and Chinese models demonstrated significant advantages in cost-effectiveness.
More from coding & agent
- Claude Code is Like Memento's Protagonist: Leaving Notes to Fight Amnesia — MattGarciaEth · 2026-08-01
- antirez Tests Domestic AI Models for Coding: DeepSeek vs GLM vs Kimi — antirez · 2026-08-01
- Coding Agents Uncover Bugs in PyTorch, vLLM and More in Just One Month — jxmnop · 2026-08-01
- AI Pipeline Ships Junk Content Despite Passing 121 Automated Tests — Positive-Emu-8379 · 2026-08-01
- 8 Types of LLMs Every GenAI Engineer Should Know for AI Agents — mdancho84 · 2026-08-01
- Autonomously Decompiling a Childhood Game via SSH While at the Park — yacineMTB · 2026-08-01