Cursor ships CursorBench 4.0; speculation swirls that Gemini 3.8 post-training differs

ivan_bezdomny · x · 2026-09-12

Cursor rolled out CursorBench 4.0, adding tasks on instruction following and long-horizon project work, with raised difficulty so all models score lower. eliebakouch speculated (unverified) that Gemini 3.8's post-training may differ notably from other models.

Related event: Cursor launches CursorBench 4.0 with harder tasks, model scores drop across the board(6 posts)→

Original post →

More from coding & agent

coding & agent channel →