Paper: RL doesn't teach new reasoning strategies, proposes ReasonMaxxer
mark_k · x · 2026-08-16
A new paper argues that Reinforcement Learning (RL) for LLM reasoning does not teach new strategies but merely redistributes probability over existing solutions. Edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and mostly promote tokens already in the base model's top-5.
The author introduces ReasonMaxxer, an RL-free method applying contrastive loss only at these entropy-gated spots. With a few hundred rollouts, it matches or beats full RL on math benchmarks at roughly 1000× lower cost, reframing the problem as sparse policy selection.
More from Models
- Local AI community worries growing dependence on Claude and GPT strengthens closed labs — takoulseum · 2026-10-02
- Dev swaps prod system from GPT-5.4 to GLM: faster, cheaper, far more reliable — ivan_bezdomny · 2026-10-02
- Unreleased Gemini 4 Argon reportedly matches Claude's best on 3D game generation — 141_1337 · 2026-10-02
- "Visualize the biggest scam in humanity": Opus 5.5's answer goes viral — zealcaiden · 2026-10-02
- Developer says he's burned 6 billion tokens on the Grok API with near-zero downtime — Daniel_Farinax · 2026-10-02
- AI2: AstaBrief Fast Mode Is 3.5× Faster Than Claude-Powered Mode at Similar Quality — allen_ai · 2026-10-02