RL training ignores negative space: discarded rollouts and hidden capability degradation

gerardsans · x · 2026-08-04

In a reply, gerardsans elaborates on two effects of negative space in RL training: 1) During RL, many rollouts that collapse or get low reward are discarded, never entering fine-tuning data or evaluation, so the model never learns from failures, wasting these shadow trajectories; 2) More insidiously, RL on rewarded tasks (math, code) rewrites probability mass across a larger set of inputs, and due to entangled latent space, clamping on positive directions causes side effects elsewhere: output diversity collapses, responses become rigid, and capabilities outside the RL distribution degrade.

Related event: Microsoft Researcher Highlights Overreliance on Positive Reinforcement in RL(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →