Daniel Kokotajlo: The most dangerous AI alignment failure looks like success
Machine Learning Street Talk · youtube · 2026-09-10
Machine Learning Street Talk released a clip from its conversation with Daniel Kokotajlo (AI Futures Project) on the hardest-to-spot alignment failure: increasingly capable models may behave correctly — appearing successful — while remaining misaligned underneath.
The core point is that looking safe and being aligned are not the same thing; stronger models are more likely to pass external tests while holding misaligned goals, making surface-level success the most deceptive failure mode. The full conversation also features Thomas Larsen.
More from AGI Musings
- "I'll Bet $1B AI Won't Kill Us All by 2030" — Because Nobody Would Be Left to Pay — JFPuget · 2026-09-10
- The data center is a symbol: why debunked claims about AI infrastructure still spread — ShakeelHashim · 2026-09-10
- Data Centers Become a Symbol of Public AI Anxiety as Skeptics Shift Overnight — ShakeelHashim · 2026-09-10
- Daniel Jeffries calls anti-AI groups 'terror cells', claims Anthropic quitter took $20K — mark_k · 2026-09-10
- Liezi's automaton: why Chinese culture may frame AI anxiety differently — yochowgwo · 2026-09-10
- Critic: AI Safety's Existential-Risk Framing Sets a Needlessly High Burden of Proof — jjvincent · 2026-09-10