Cursor’s Composer 2.5 beats Kimi on coding but loses badly on math and science

gleech · x · 2026-07-21

A new write-up tests whether post-training can specialize a general model without preserving broad reasoning.

The reported result is stark: Cursor’s Composer 2.5, a coding finetune of Kimi, improves strongly on visual reasoning, RPG-style games, and agentic coding, but loses heavily on math and scientific reasoning benchmarks. The thread argues this is evidence against the old assumption that coding training automatically improves general reasoning.

Original post →

More from coding & agent

coding & agent channel →