CursorBench ranks coding agents: Fable 5.1 Max tops at 73.4%, Grok 4.6 best value

mattyp · x · 2026-09-03

Cursor released CursorBench 3.2, evaluating coding agents on ambiguous multi-file tasks from real Cursor sessions. Key takeaways:

Related event: Cursor Launches CursorBench 3.2: Fable 5.1 Tops Coding Agent Leaderboard(4 posts)→

Original post →

More from Models

Models channel →