Base Labs: RL gains mostly come from easier problems and fixing doom loops

baseten · x · 2026-10-06

Base Labs shares weekly 'rollouts' of ongoing work. This week they probed the low-dimensional structure of RL training and tested whether recent findings on rank-1 extrapolation of early updates hold. They find performance gains largely come from easier problems and fixing doom loops, while harder environments are less linearly extrapolatable.

Original post →

More from Research

Research channel →