Paper: RL for LLM Reasoning is Sparse Policy Selection, Not Capability Learning
mark_k · x · 2026-08-16
A new paper analyzes reinforcement learning (RL) for LLM reasoning and finds that it does not teach new strategies; instead, it merely redistributes probability over solutions the base model already contains. The beneficial edits are sparse (1–3% of tokens), concentrated at high-entropy decision points, and almost always promote tokens already in the base model’s top-5.
Building on this, the authors introduce ReasonMaxxer, an RL-free method that applies contrastive loss only at those entropy-gated spots. It requires only a few hundred rollouts, tens of problems, and minutes on a single GPU. Despite costing 1000× less, ReasonMaxxer matches or beats full RL across models and math benchmarks, reframing the problem as sparse policy selection rather than capability acquisition.
More from Research
- Molecular Glue Reprograms BCL6, Clears Aggressive Tumors in Mice in 11 Days — WmHaseltine · 2026-10-03
- Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive — dlwh · 2026-10-03
- Harvard-led team unveils brain imaging 60x faster than fMRI, tracking activity in ~100ms — melnykowycz · 2026-10-02
- RExBench: best coding agent implements research extensions only 33% of the time — najoungkim · 2026-10-02
- Retrieval-augmented episodic memory narrows LLMs' lexical frequency gap in syntax tests — najoungkim · 2026-10-02
- EMPIRIC teaches robots missing physics as code, solving all 25 tasks where baselines get 14-16 — tomssilver · 2026-10-02