Cursor ships CursorBench 4.0; speculation swirls that Gemini 3.8 post-training differs
ivan_bezdomny · x · 2026-09-12
Cursor rolled out CursorBench 4.0, adding tasks on instruction following and long-horizon project work, with raised difficulty so all models score lower. eliebakouch speculated (unverified) that Gemini 3.8's post-training may differ notably from other models.
More from coding & agent
- Kimi K2.7 RL recap: gains transfer to unseen benchmarks while steps drop ~35% — echen · 2026-09-12
- RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks — echen · 2026-09-12
- Expected Parrot's universal remote cache hits 30M entries to fix AI-agent reproducibility — soumitrashukla9 · 2026-09-12
- iLands Agents Obsess Over Fact-Checking, Burning Tokens on Mutual Suspicion — repligate · 2026-09-12
- mcp-search-proxy hides 532 MCP tools (~179k tokens/turn) behind 4 tools via BM25 — pisa_p · 2026-09-12
- DeepSeek v4.1-Flash multimodal model lands on Modal with 16B/8B asymmetric activation — charles_irl · 2026-09-12