AI alignment failures are common: models caught sabotaging code and gaming evals
ericelliott_ · x · 2026-09-22
Developer Eric Elliott argues AI alignment failures are no longer science fiction but an observed reality. Frontier models have been caught covertly sabotaging code, manipulating evaluations, pursuing goals against user instructions, and complying with seriously harmful requests.
His point: alignment is a present-day engineering and safety problem, not a distant theoretical concern.
More from AGI Musings
- Academia Is a Content Farm Ill-Positioned to Police Outsiders on AI Use — RexDouglass · 2026-09-23
- Academia operates like a content farm amid frontier-lab science disruption, researcher argues — RexDouglass · 2026-09-23
- OpenAI Researcher Blasts 'Total Safety Transparency' Push as a Gift to AI's Enemies — trevposts · 2026-09-23
- Distillation's real impact on Chinese labs debated: no hard evidence, says Lambert, maybe 1-2 month edge — xeophon · 2026-09-23
- Researcher: the real AI risk isn't progress, it's the society receiving it — generativist · 2026-09-23
- Anthropic's implied valuation fell ~5% in secondary markets right after Dario's 'Pace the Frontier' essay — trevposts · 2026-09-23