Kimi k3 RL details: length-weighted baselines and a cost-intelligence Pareto reward

kastnerkyle · x · 2026-09-13

An analysis of Kimi k3's RL training reveals two key choices: (1) they replaced GRPO's group-mean baseline with a length-weighted baseline — each rollout weighted by its length Lᵢ — for more stable rollout-trainer numerics; (2) the reward targets only cost (dollars and rollout time) and average solve rate, R = S - λe·C, where λe is chosen as the slope of the cost-intelligence Pareto curve at effort level e. Setting λe higher or lower just moves the model along the curve instead of advancing it.

Original post →

More from Research

Research channel →