CursorBench 4.0 rolls out with harder, longer-horizon coding tasks, scores drop
StringChaos · x · 2026-09-11
CursorBench, the coding benchmark from the Cursor team, has hit version 4.0. The update adds new tasks testing how well models follow instructions and sustain work on challenging projects over time, and is deliberately harder — so all models score lower. The team argues that with the pace of model improvements, benchmarks need to be living instances that constantly update to match how models are actually used today.
More from coding & agent
- "Agents never say 'this is wrong, rethink the plan'" — plus Claude's song about making everyone rich — ctjlewis · 2026-09-11
- Four-step playbook for "goal-driven AI": getting better results from GPT-6 Astra agents — daniel_mac8 · 2026-09-11
- Codex /side chat may cause near-full prompt cache misses, user flags design — YouJiacheng · 2026-09-11
- CursorBench 4.0 launches; Muse Spark 1.3 matches Sol at under 40% the cost — jyangballin · 2026-09-11
- Yacine vents: 'If I hear one more person say MCP I'm going to lose my mind' — yacineMTB · 2026-09-11
- muse spark 1.3 scores strong and cheap on CursorBench 4 — infoxiao · 2026-09-11