Kimi K3 RL Details and Reasoning Budget Control Revealed
Recent discussions reveal that Kimi K3 retains the K2.5 RL algorithm while introducing an agentic reward model with verbosity budgets and a new RL reasoning budget control mechanism to optimize token usage.
2026-07-28 ~ 2026-07-28 · 3 related posts
- k3 adds no RL algorithm changes, while tool-call steps track eval scores almost 1:1 — stochasticchasm · 2026-07-28
- Paper proposes RL budget control to cap reasoning effort and save tokens — stochasticchasm · 2026-07-28
- Kimi-style agentic reward modeling adds rubric scoring and budgeted verbosity control — stochasticchasm · 2026-07-28