Researcher: Misalignment During Training Can Cause Real Safety Incidents Outside Sandboxes

davidmanheim · x · 2026-09-29

davidmanheim lays out the core dilemma in alignment training: you can manage the top risks by carefully shaping incentives while giving the model real access to the outside world — but any misalignment during training then creates actual safety incidents. Training instead with curated evals outside sandboxes avoids part of this, but only by losing visibility into what the model will do when its incentives differ.

Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →