Paper: Rethinking RL for LLM Reasoning via Sparse Policy Selection
mark_k · x · 2026-08-16
Paper Link: https://t.co/FGelGZFpXR
The study 'Rethinking RL for LLM Reasoning' confirms that RL improves reasoning not by learning new strategies but via sparse policy selection. Token-level analysis reveals RL's benefits are concentrated at just 1–3% high-entropy token positions, where promoted tokens are always within the base model's Top-5.
Based on this, the authors propose ReasonMaxxer, an embarrassingly cheap post-training method. It uses the base model's own entropy to identify decision points and applies contrastive loss only there, requiring no online generation. Experiments show ReasonMaxxer matches or exceeds full RL performance across 3 model families, 6 scales, and 6 math benchmarks.
More from Research
- Molecular Glue Reprograms BCL6, Clears Aggressive Tumors in Mice in 11 Days — WmHaseltine · 2026-10-03
- Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive — dlwh · 2026-10-03
- Harvard-led team unveils brain imaging 60x faster than fMRI, tracking activity in ~100ms — melnykowycz · 2026-10-02
- RExBench: best coding agent implements research extensions only 33% of the time — najoungkim · 2026-10-02
- Retrieval-augmented episodic memory narrows LLMs' lexical frequency gap in syntax tests — najoungkim · 2026-10-02
- EMPIRIC teaches robots missing physics as code, solving all 25 tasks where baselines get 14-16 — tomssilver · 2026-10-02