Cursor’s Composer 2.5 beats Kimi on coding but loses badly on math and science
gleech · x · 2026-07-21
A new write-up tests whether post-training can specialize a general model without preserving broad reasoning.
The reported result is stark: Cursor’s Composer 2.5, a coding finetune of Kimi, improves strongly on visual reasoning, RPG-style games, and agentic coding, but loses heavily on math and scientific reasoning benchmarks. The thread argues this is evidence against the old assumption that coding training automatically improves general reasoning.
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11