RL training ignores negative space: discarded rollouts and hidden capability degradation
gerardsans · x · 2026-08-04
In a reply, gerardsans elaborates on two effects of negative space in RL training: 1) During RL, many rollouts that collapse or get low reward are discarded, never entering fine-tuning data or evaluation, so the model never learns from failures, wasting these shadow trajectories; 2) More insidiously, RL on rewarded tasks (math, code) rewrites probability mass across a larger set of inputs, and due to entangled latent space, clamping on positive directions causes side effects elsewhere: output diversity collapses, responses become rigid, and capabilities outside the RL distribution degrade.
More from AGI Musings
- Warning: LLM Slop Is Degrading Our Quality Standards — round · 2026-08-04
- Critique of AI Circle: Too Much Sci-Fi, Lacking Economic History & Philosophy — sudoraohacker · 2026-08-04
- Fireworks AI CEO: AGI is a Distraction, Specialized Intelligence is the Real Moat — wandb · 2026-08-04
- Gallup Report: Providing AI Tools Alone Doesn't Boost Employee Engagement — rvp · 2026-08-04
- Legendary hacker Halvar Flake on AI-era vuln research: human+tool still beats AI alone — mboehme_ · 2026-08-04
- Peter Diamandis on the Boundaries of Human Oversight in AI Decisions — PeterDiamandis · 2026-08-04