Cursor’s Composer 2.5 looks much worse at reasoning than its Kimi base model
gleech · x · 2026-07-23
A quoted thread says a set of launch posts points to a hunch: post-training may have sharp limits.
In experiments led by @ncznc and @peligrietzer, Cursor’s Composer 2.5 looks very different from its Kimi base model and is much worse at reasoning. The reaction from another researcher is that this is a useful natural experiment against the idea that strong RL generalization is easy, even if scaling can still paper over some practical gaps.
More from coding & agent
- Users say Linear’s Slack agent still lacks memory, speed and thread context — var_epsilon · 2026-07-23
- A simple quant benchmark could expose frontier-model failures fast — PtrPomorski · 2026-07-23
- Codex Security plugin returns as an open-source codebase scanner with fix generation — reach_vb · 2026-07-23
- Claude Code v2.1.218 moves code review to a background subagent — ashwin-ant · 2026-07-23
- Ninja pitches a subnet that rewards agents for doing more with fewer tokens — const_reborn · 2026-07-23
- Bittensor adds miner collateral to subnet registration and releases it with emissions — const_reborn · 2026-07-23