Cursor’s Composer 2.5 beats Kimi on coding but loses badly on math and science
gleech · x · 2026-07-21
A new write-up tests whether post-training can specialize a general model without preserving broad reasoning.
The reported result is stark: Cursor’s Composer 2.5, a coding finetune of Kimi, improves strongly on visual reasoning, RPG-style games, and agentic coding, but loses heavily on math and scientific reasoning benchmarks. The thread argues this is evidence against the old assumption that coding training automatically improves general reasoning.
More from coding & agent
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11