Kimi k3 RL details: length-weighted baselines and a cost-intelligence Pareto reward
kastnerkyle · x · 2026-09-13
An analysis of Kimi k3's RL training reveals two key choices: (1) they replaced GRPO's group-mean baseline with a length-weighted baseline — each rollout weighted by its length Lᵢ — for more stable rollout-trainer numerics; (2) the reward targets only cost (dollars and rollout time) and average solve rate, R = S - λe·C, where λe is chosen as the slope of the cost-intelligence Pareto curve at effort level e. Setting λe higher or lower just moves the model along the curve instead of advancing it.
More from Research
- Flash-BoN at ECCV 2026: rethinking diffusion inference-time scaling with wall-clock time as the budget — RisingSayak · 2026-09-13
- Biologists push back: AlphaFold useful in some cases but weak on out-of-distribution proteins — lpachter · 2026-09-13
- IR Papers Weekly Vol.173: Spectral Compression for ColBERT, MoE Multimodal Embeddings from Meta — _reachsumit · 2026-09-13
- Astra and Fable: Simple 2025-Style Alignment Evals Still Get Hacked — Levitating · 2026-09-13
- TRL v1.13 ships long-context training: 1M+ token sequences on a single 8-GPU H100 node — SergioPaniego · 2026-09-13
- Robotics researcher Yacine mocks LLM-inspired RL robotics papers as 'basically wrong' — yacineMTB · 2026-09-13