RL mostly sharpens easy tasks, but can find tail solutions on harder ones

Pavel_Izmailov · x · 2026-07-21

The authors examine what RL actually changes in the policy.

Related event: New Research Proposes Joint Scaling Law for Pretraining and RL(18 posts)→

Original post →

More from Research

Research channel →