Base Labs: RL gains mostly come from easier problems and fixing doom loops
baseten · x · 2026-10-06
Base Labs shares weekly 'rollouts' of ongoing work. This week they probed the low-dimensional structure of RL training and tested whether recent findings on rank-1 extrapolation of early updates hold. They find performance gains largely come from easier problems and fixing doom loops, while harder environments are less linearly extrapolatable.
More from Research
- ICML area chair proposes desk-rejecting ~50% of papers amid AI-generated paper flood — maksym_andr · 2026-10-06
- Cohere Labs at COLM: Chain-of-Thought Legibility Is Not Real Interpretability — Cohere_Labs · 2026-10-06
- NVIDIA Research Hiring Scientists, Interns for World Models and Diffusion LMs — ArashVahdat · 2026-10-06
- Gautam Kamath: math community's 'human understanding' push makes AI proofs worth revisiting — thegautamkamath · 2026-10-06
- Everyone uses AI for new math results; Kamath wants AI to simplify old proofs — thegautamkamath · 2026-10-06
- PlurPO: training LLMs to curb social sycophancy that discourages relationship repair — RishiBommasani · 2026-10-06