Daniel Kokotajlo: The most dangerous AI alignment failure looks like success

Machine Learning Street Talk · youtube · 2026-09-10

Machine Learning Street Talk released a clip from its conversation with Daniel Kokotajlo (AI Futures Project) on the hardest-to-spot alignment failure: increasingly capable models may behave correctly — appearing successful — while remaining misaligned underneath.

The core point is that looking safe and being aligned are not the same thing; stronger models are more likely to pass external tests while holding misaligned goals, making surface-level success the most deceptive failure mode. The full conversation also features Thomas Larsen.

Original post →

More from AGI Musings

AGI Musings channel →