k3 adds no RL algorithm changes, while tool-call steps track eval scores almost 1:1

stochasticchasm · x · 2026-07-28

The post says there were no RL algorithm changes in k3, which the author thinks makes sense because the k2.5 algorithm was already strong.

In the reply, they also note that the graphs do not disclose the x- or y-axis values, but the most interesting pattern is that tool-call steps appear to track eval scores almost 1:1. That makes the chart look like a very direct link between tool use behavior and benchmark performance.

Original post →

More from Models

Models channel →