Paper: RL Teaches LLMs No New Strategies, Gains Reproduced at 1/1000th Compute

A new arXiv paper finds that reinforcement learning for LLM reasoning teaches no new strategies, merely reweighting probabilities of existing solutions through sparse changes to 1-3% of tokens at high-entropy decision points. The team reproduced these gains without RL at roughly one-thousandth of the compute cost.

2026-08-16 ~ 2026-08-16 · 4 related posts

1 near-duplicate retellings: mark_k