Paper claims RL for reasoning changes only 1-3% of tokens, gains replicated without RL

juanviera23 · reddit · 2026-08-16

An arXiv paper claims that Reinforcement Learning for reasoning alters only 1-3% of tokens. The researchers replicated the performance gains without RL, using approximately 1000x less compute, raising questions about the necessity of high-cost reasoning training.

Related event: Paper: RL Teaches LLMs No New Strategies, Gains Reproduced at 1/1000th Compute(4 posts)→

Original post →

More from Research

Research channel →