Microsoft Researcher Highlights Overreliance on Positive Reinforcement in RL

Microsoft Principal Researcher Alexia Jolicoeur discusses the limitations of focusing solely on positive reinforcement in RL training. Discarding low-reward rollouts prevents models from learning from failures, potentially leading to capability degradation.

2026-08-04 ~ 2026-08-04 · 2 related posts