Paper: RL Teaches LLMs No New Strategies, Gains Reproduced at 1/1000th Compute
A new arXiv paper finds that reinforcement learning for LLM reasoning teaches no new strategies, merely reweighting probabilities of existing solutions through sparse changes to 1-3% of tokens at high-entropy decision points. The team reproduced these gains without RL at roughly one-thousandth of the compute cost.
2026-08-16 ~ 2026-08-16 · 4 related posts
- Paper claims RL for reasoning changes only 1-3% of tokens, gains replicated without RL — juanviera23 · 2026-08-16
- Paper: RL doesn't teach new reasoning strategies, proposes ReasonMaxxer — mark_k · 2026-08-16
- Paper: Rethinking RL for LLM Reasoning via Sparse Policy Selection — mark_k · 2026-08-16
1 near-duplicate retellings: mark_k