RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks
echen · x · 2026-09-12
Post-training Kimi K2.7 (Max reasoning) with pure RL on just 1,700 Surge coding tasks improved it across all five external benchmarks: +20.0 SWE-Marathon, +14.6 Terminal-Bench 2.1, +12.4 DeepSWE, +10.7 Terminal-Bench 3, +4.7 SWE-Bench Pro.
The framing: before RL, K2.7 could write code like an intern but failed the last mile—dropping requirements, writing narrow tests, creating regressions. After RL it ships like a staff engineer.
Related event: RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks(2 posts)→
More from coding & agent
- Cloud-computer agent fails to log into Costco and order, spawning a 'dad eval bench' idea — msg · 2026-09-12
- New morning routine: reviewing overnight agent work from bed via phone — evielync · 2026-09-12
- Anthropic team uses Claude for on-call: alert triage, root cause and fix proposals — ClaudeDevs · 2026-09-12
- Developer streams an AI agent playing Minecraft all day via Codex, aiming to build an auto chicken farm — nickbaumann_ · 2026-09-12
- Why working with AI agents all day is exhausting: delegation adds mental load, not removes it — bendee983 · 2026-09-12
- "Codex is my main interface to the system": user says agent made Linux admin effortless — mark_k · 2026-09-12