Microsoft Researcher Highlights Overreliance on Positive Reinforcement in RL
Microsoft Principal Researcher Alexia Jolicoeur discusses the limitations of focusing solely on positive reinforcement in RL training. Discarding low-reward rollouts prevents models from learning from failures, potentially leading to capability degradation.
2026-08-04 ~ 2026-08-04 · 2 related posts
- Microsoft Researcher Discusses RL: Why Does the Industry Only Focus on Positive Reinforcement? — gerardsans · 2026-08-04
- RL training ignores negative space: discarded rollouts and hidden capability degradation — gerardsans · 2026-08-04