Kimi K2.7 RL recap: gains transfer to unseen benchmarks while steps drop ~35%

echen · x · 2026-09-12

Surge AI's detailed recap of RL post-training on Kimi K2.7 (Max reasoning) reports three key findings:

Core conclusion: K2.7 already knew how to code—RL made its execution less brittle, turning intern-style code into shippable code. The team studied trajectories to identify the specific new behaviors learned.

Related event: RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →