Microsoft Researcher Discusses RL: Why Does the Industry Only Focus on Positive Reinforcement?
gerardsans · x · 2026-08-04
In an academic talk show, the host invited Alexia Jolicoeur-Martineau, Principal Researcher at Microsoft and author of the Tiny Recursive Model (which scored 45% on ARC-AGI-1 and won an award), to discuss Reinforcement Learning (RL).
Addressing a community question, the host raised a critical point: the current AI industry applying RL tends to focus only on the performance gains from positive reinforcement, largely ignoring the negative effects of rollouts that collapse during RL training and are excluded from the fine-tuning dataset or evaluation process.
More from Research
- Nature Study: AI Dermatology Diagnosis Amplifies Public Automation Bias — EricTopol · 2026-08-04
- Radical Co-founder on Training the Largest Genome Model to Write DNA — exnx · 2026-08-04
- 3 Lines of Code Fixed 123 Failed PPO Experiments by Changing Reward Shaping — mikeysce · 2026-08-04
- Research: Frozen Pixel-Space Diffusion Models Can Self-Guide — nanyang-technological-university-singapore · 2026-08-04
- ICDAR 2026 Competition: Multimodal AI Struggles with Scientific Figures — SciKnowOrg · 2026-08-04
- GEOID-Flood: Large-Scale Multi-Modal Benchmark for Flood Segmentation — links-ads · 2026-08-04