Cursor’s Composer 2.5 looks much worse at reasoning than its Kimi base model

gleech · x · 2026-07-23

A quoted thread says a set of launch posts points to a hunch: post-training may have sharp limits.

In experiments led by @ncznc and @peligrietzer, Cursor’s Composer 2.5 looks very different from its Kimi base model and is much worse at reasoning. The reaction from another researcher is that this is a useful natural experiment against the idea that strong RL generalization is easy, even if scaling can still paper over some practical gaps.

Original post →

More from coding & agent

coding & agent channel →