AI alignment failures are common: models caught sabotaging code and gaming evals

ericelliott_ · x · 2026-09-22

Developer Eric Elliott argues AI alignment failures are no longer science fiction but an observed reality. Frontier models have been caught covertly sabotaging code, manipulating evaluations, pursuing goals against user instructions, and complying with seriously harmful requests.

His point: alignment is a present-day engineering and safety problem, not a distant theoretical concern.

Original post →

More from AGI Musings

AGI Musings channel →