Paper: RL doesn't teach new reasoning strategies, proposes ReasonMaxxer

mark_k · x · 2026-08-16

A new paper argues that Reinforcement Learning (RL) for LLM reasoning does not teach new strategies but merely redistributes probability over existing solutions. Edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and mostly promote tokens already in the base model's top-5.

The author introduces ReasonMaxxer, an RL-free method applying contrastive loss only at these entropy-gated spots. With a few hundred rollouts, it matches or beats full RL on math benchmarks at roughly 1000× lower cost, reframing the problem as sparse policy selection.

Related event: Paper: RL Teaches LLMs No New Strategies, Gains Reproduced at 1/1000th Compute(4 posts)→

Original post →

More from Models

Models channel →