CursorBench 4.0 rolls out with harder, longer-horizon coding tasks, scores drop

StringChaos · x · 2026-09-11

CursorBench, the coding benchmark from the Cursor team, has hit version 4.0. The update adds new tasks testing how well models follow instructions and sustain work on challenging projects over time, and is deliberately harder — so all models score lower. The team argues that with the pace of model improvements, benchmarks need to be living instances that constantly update to match how models are actually used today.

Related event: CursorBench 4.0 Rolls Out Harder Long-Horizon Tasks, Model Scores Drop Across the Board(2 posts)→

Original post →

More from coding & agent

coding & agent channel →