Cursor’s Composer 2.5 beats Kimi on coding but loses badly on math and science
gleech · x · 2026-07-21
A new write-up tests whether post-training can specialize a general model without preserving broad reasoning.
The reported result is stark: Cursor’s Composer 2.5, a coding finetune of Kimi, improves strongly on visual reasoning, RPG-style games, and agentic coding, but loses heavily on math and scientific reasoning benchmarks. The thread argues this is evidence against the old assumption that coding training automatically improves general reasoning.
More from coding & agent
- A Forward Deployed Engineer job really has three stages: audit, evals, deploy — blaizedsouza · 2026-07-22
- 438 sealed tests show coding agents prefer DIY over third-party databases — cramforce · 2026-07-22
- Tenable and AWS launch a Black Hat build event for open-source security agents and MCP servers — Dave_Maynor · 2026-07-22
- Codex helps build Valdiluce, an open-world game with climbing, gliding and gondolas — Dimillian · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22