RL Post-Training on 1,700 Tasks Boosts Kimi K2.7 Across Coding Benchmarks

Surge AI post-trained Kimi K2.7 (Max reasoning) with pure RL on 1,700 self-built coding tasks, improving it across five external coding benchmarks. Gains generalized to unseen evaluations while trajectory steps dropped by about 30%.

2026-09-12 ~ 2026-09-12 · 2 related posts