RL-Only Post-Training Lifts Kimi K2.7 Past GPT-5.6 and Kimi K3 on Coding Benchmarks
echen · x · 2026-09-17
Surge AI post-trained Kimi K2.7 (Max reasoning) with reinforcement learning alone on 1,700 coding tasks and improved scores on all five external benchmarks. Gains transferred across three agent harnesses and to benchmarks that didn't exist at training time; the smaller post-trained model beat Kimi's larger K3 on Terminal-Bench 2.1/3 and slightly edged GPT-5.6 Sol on SWE-Bench Pro. Median trajectory length dropped from 150 to 98 steps on DeepSWE. A favorite emergent example: with no zstd binary available to verify a Zstandard decompressor, the model wrote its own compressor first, generated valid test files, and used them to test its work.
More from coding & agent
- OpenClaw creator to join Cloudflare Connect 2026 panel on agentic platforms — steipete · 2026-09-17
- Your LLM doesn't understand MCP — keeping the tool-call boundary clear makes agents easier to debug — gethackteam · 2026-09-17
- YC-backed Extend launches Parse Router to route each page to the right parsing engine — ycombinator · 2026-09-17
- He had Codex make a phone call to redeem a $500 gift card — and it worked — brandon_galang · 2026-09-17
- Alchemy ships official coding-agent prompt for its TypeScript IaC tool — samgoodwin89 · 2026-09-17
- LangChain founder joins Navigators podcast: outcome-first agent building and model gateways as core infra — LangChain · 2026-09-17