Kimi K2.7 RL recap: gains transfer to unseen benchmarks while steps drop ~35%
echen · x · 2026-09-12
Surge AI's detailed recap of RL post-training on Kimi K2.7 (Max reasoning) reports three key findings:
- Gains transfer: training tasks predate DeepSWE, SWE-Marathon, and Terminal-Bench 3; improvements hold across three different agent harnesses with different tools, prompting, and loop structures. SWE-Marathon requires multi-hour system-building trajectories; TB3 spans hardware, scientific computing, systems, and security.
- Smaller model beats larger ones: per K3's self-reported scores, the post-trained K2.7 leads K3 on Terminal-Bench 2.1 and 3, and slightly edges GPT-5.6 Sol on SWE-Bench Pro (contextual, not controlled comparisons).
- More efficient: median trajectory length fell from 150 to 98 steps on DeepSWE and 102 to 78 on Terminal-Bench 3.
Core conclusion: K2.7 already knew how to code—RL made its execution less brittle, turning intern-style code into shippable code. The team studied trajectories to identify the specific new behaviors learned.
Related event: RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks(2 posts)→
More from coding & agent
- Cloud-computer agent fails to log into Costco and order, spawning a 'dad eval bench' idea — msg · 2026-09-12
- Anthropic team uses Claude for on-call: alert triage, root cause and fix proposals — ClaudeDevs · 2026-09-12
- Developer streams an AI agent playing Minecraft all day via Codex, aiming to build an auto chicken farm — nickbaumann_ · 2026-09-12
- Why working with AI agents all day is exhausting: delegation adds mental load, not removes it — bendee983 · 2026-09-12
- "Codex is my main interface to the system": user says agent made Linux admin effortless — mark_k · 2026-09-12
- Alchemy's distilled project builds unified agent-friendly SDKs for 60+ SaaS services — samgoodwin89 · 2026-09-12