RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks

echen · x · 2026-09-12

Post-training Kimi K2.7 (Max reasoning) with pure RL on just 1,700 Surge coding tasks improved it across all five external benchmarks: +20.0 SWE-Marathon, +14.6 Terminal-Bench 2.1, +12.4 DeepSWE, +10.7 Terminal-Bench 3, +4.7 SWE-Bench Pro.

The framing: before RL, K2.7 could write code like an intern but failed the last mile—dropping requirements, writing narrow tests, creating regressions. After RL it ships like a staff engineer.

Related event: RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →