CursorBench ranks coding agents: Fable 5.1 Max tops at 73.4%, Grok 4.6 best value
mattyp · x · 2026-09-03
Cursor released CursorBench 3.2, evaluating coding agents on ambiguous multi-file tasks from real Cursor sessions. Key takeaways:
- Top scores: Fable 5.1 Max leads at 73.4% ($9.64/task); Fable 5.1 Extra High follows at 72.8% ($6.95/task).
- Value picks: Grok 4.6 Extra High hits 70.8% at just $2.81/task; GPT-5.6 Luna Max scores 61.1% at $0.39/task.
- Premium tier: Fable 5 Max scores 70.5% but costs $17.32/task; Opus 5 Max hits 70.0% at $8.23.
- The board includes token usage and step counts for cost-capability-efficiency comparison.
Related event: Cursor Launches CursorBench 3.2: Fable 5.1 Tops Coding Agent Leaderboard(4 posts)→
More from Models
- Baseten ships GLM-5.3 Fast: speed-optimized open-weight model for real-time workloads — baseten · 2026-09-03
- Users Report Claude Racking Up Daily Mistakes and Hallucinations — lilyraynyc · 2026-09-03
- First run of Gemini 3.8 Flash fails: model keeps thinking until it times out — rickasaurus · 2026-09-03
- Early user verdict: fable 5 outperforms the newer fable 5.1 — BLUECOW009 · 2026-09-03
- Flash 3.8 Review: Great When Working, but Stuck in Silent Token-Burning Loops — brandon_galang · 2026-09-03
- Claude Max 20x buyer says weekly limits, not the 5-hour window, are the real bottleneck; r/ClaudeAI deleted his post — conorearly · 2026-09-03