k3 adds no RL algorithm changes, while tool-call steps track eval scores almost 1:1
stochasticchasm · x · 2026-07-28
The post says there were no RL algorithm changes in k3, which the author thinks makes sense because the k2.5 algorithm was already strong.
In the reply, they also note that the graphs do not disclose the x- or y-axis values, but the most interesting pattern is that tool-call steps appear to track eval scores almost 1:1. That makes the chart look like a very direct link between tool use behavior and benchmark performance.
More from Models
- Kimi K3 launches on Together AI with Day 0 access for coding agents — togethercompute · 2026-07-28
- Kimi K3 throughput jumps from 19 to 49 tok/s on OpenRouter — cedric_chee · 2026-07-28
- Kimi K3 launches with 2.8T parameters, 1M context and $0.30 input pricing — togethercompute · 2026-07-28
- Together AI brings Moonshot’s Kimi K3 online with 1M context and agent tools — togethercompute · 2026-07-28
- DeepInfra adds Claude Opus 5 with 1M context and $5/$25 pricing — gharik · 2026-07-28
- Kimi K3 goes live on OpenRouter as third-party providers race to add support — scaling01 · 2026-07-28