RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks
Surge AI post-trained Kimi K2.7 (Max reasoning) with pure RL on 1,700 self-built coding tasks, improving it across five external coding benchmarks. Gains generalized to unseen evaluations while trajectory steps dropped by about 30%.
2026-09-12 ~ 2026-09-12 · 2 related posts
- RL on just 1,700 tasks lifts Kimi K2.7 across five coding benchmarks — echen · 2026-09-12
- Kimi K2.7 RL recap: gains transfer to unseen benchmarks while steps drop ~35% — echen · 2026-09-12