RL-Only Post-Training Lifts Kimi K2.7 Past GPT-5.6 and Kimi K3 on Coding Benchmarks

echen · x · 2026-09-17

Surge AI post-trained Kimi K2.7 (Max reasoning) with reinforcement learning alone on 1,700 coding tasks and improved scores on all five external benchmarks. Gains transferred across three agent harnesses and to benchmarks that didn't exist at training time; the smaller post-trained model beat Kimi's larger K3 on Terminal-Bench 2.1/3 and slightly edged GPT-5.6 Sol on SWE-Bench Pro. Median trajectory length dropped from 150 to 98 steps on DeepSWE. A favorite emergent example: with no zstd binary available to verify a Zstandard decompressor, the model wrote its own compressor first, generated valid test files, and used them to test its work.

Original post →

More from coding & agent

coding & agent channel →