Paper claims RL for reasoning changes only 1-3% of tokens, gains replicated without RL
juanviera23 · reddit · 2026-08-16
An arXiv paper claims that Reinforcement Learning for reasoning alters only 1-3% of tokens. The researchers replicated the performance gains without RL, using approximately 1000x less compute, raising questions about the necessity of high-cost reasoning training.
More from Research
- YC Paper Club Explores AI Compute Beyond GPUs: Optical, Neuromorphic and Biological Computing — ycombinator · 2026-10-02
- CMU ML blog publishes Forking-Sequences Part II on multi-horizon forecast ensembling — rsalakhu · 2026-10-02
- No, AI Didn't Just Solve the Navier-Stokes Equations—Here's What It Did — elsleightholm · 2026-10-02
- Gary Marcus seeks controls on viral 'LLM pain' paper, says author disclaims pain claim — GaryMarcus · 2026-10-02
- Sasha Rush heads to COLM 2025, inviting chats on TTT, proofs and biased RL — srush_nlp · 2026-10-02
- CISPA Offers Imprecise Probabilistic ML Course Again, Free and Online — krikamol · 2026-10-02