Paper: RL doesn't teach new reasoning strategies, proposes ReasonMaxxer
mark_k · x · 2026-08-16
A new paper argues that Reinforcement Learning (RL) for LLM reasoning does not teach new strategies but merely redistributes probability over existing solutions. Edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and mostly promote tokens already in the base model's top-5.
The author introduces ReasonMaxxer, an RL-free method applying contrastive loss only at these entropy-gated spots. With a few hundred rollouts, it matches or beats full RL on math benchmarks at roughly 1000× lower cost, reframing the problem as sparse policy selection.
More from Models
- Zero-Dataset Fine-tuning: Arthemy Suite Reshapes Krea-2 in Real-Time — ItalianArtProfessor · 2026-08-17
- Open Source Frontier Lag Shrinks: 30B Models by 2027 — PetersOdyssey · 2026-08-17
- RL investments pay off; Laguna S2.2 expected to lead open models — andrew_n_carr · 2026-08-17
- Blackbox AI to Update GLM 5.2 on Monday, Potentially Doubling Speed — brandon_galang · 2026-08-17
- Opus 5 Ultracode costs $5 just to explore codebase — haydendevs · 2026-08-17
- User Test: Minimal Difference Between Google AI 3.6 and 3.7 Flash — gaganghotra_ · 2026-08-17