Paper claims RL for reasoning changes only 1-3% of tokens, gains replicated without RL
juanviera23 · reddit · 2026-08-16
An arXiv paper claims that Reinforcement Learning for reasoning alters only 1-3% of tokens. The researchers replicated the performance gains without RL, using approximately 1000x less compute, raising questions about the necessity of high-cost reasoning training.
More from Research
- RL may train generalized dispositions, with model values shaped by training environment structure — novasarc01 · 2026-08-17
- Papers should ship with data and code: AI-authorship panic vanishes with verification — ipeirotis · 2026-08-17
- Deepest Layer Suboptimal for Protein Language Models — anshulkundaje · 2026-08-17
- AI benchmarks in non-verifiable domains should adopt qualitative research methodology — emollick · 2026-08-17
- Paper claims the third pretraining axis is freedom, not exploration — teortaxesTex · 2026-08-17
- IJCAI 2026 Day 1: Cognitive Robotics and Graph Distribution Shifts — banazir · 2026-08-17