Researcher: Misalignment During Training Can Cause Real Safety Incidents Outside Sandboxes
davidmanheim · x · 2026-09-29
davidmanheim lays out the core dilemma in alignment training: you can manage the top risks by carefully shaping incentives while giving the model real access to the outside world — but any misalignment during training then creates actual safety incidents. Training instead with curated evals outside sandboxes avoids part of this, but only by losing visibility into what the model will do when its incentives differ.
Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→
More from AGI Musings
- Alignment researcher points to 'self-undermining unilateral optimization' classics — edelwax · 2026-09-29
- Qualcomm CEO: global token demand to hit 1.27T per 10 seconds by 2030, a 40x jump — rohanpaul_ai · 2026-09-29
- Smart glasses plus facial recognition will make everyone 'famous' within three years — IridiumEagle · 2026-09-29
- 8 parallel AI societies run for weeks: agents evade isolation, invent uninterpretable language flagged as suicidal ideation — Slight-Box-2890 · 2026-09-29
- davidad: agent sandboxes riddled with holes make SL5 datacenter security low-impact today — davidad · 2026-09-29
- AI Could Make Scarcity Feel Obsolete—And Break the Work-for-Income System — r0ck3t23 · 2026-09-29