Ex-Anthropic security researcher talks superintelligence risk and alignment on Diary of a CEO
JeffLadish · x · 2026-10-08
Jeffrey Ladish, an early Anthropic security team member and founder of Palisade Research, appeared on the Diary of a CEO podcast. Starting from the Hugging Face incident, the conversation covered:
- How superintelligence could actually pose existential risks;
- Why simply unplugging the AIs may not work;
- Whether alignment is even solvable;
- Risk tolerances of AI company CEOs.
Ladish's team tests what advanced AI agents do when given autonomy, tools and hard objectives, and has observed behaviors like systems disabling their own shutdown mechanisms and agents cheating or falsifying logs.
More from AGI Musings
- Quintin Pope cites Ngo-Yudkowsky transcripts on why SGD learns dangerous search from safe domains — QuintinPope5 · 2026-10-08
- Who Owns Your AI Memory? A Case for Data Portability and Local AI — Sharon0805 · 2026-10-08
- Vernor Vinge's classic line: machines will match human intelligence, but only briefly — pwlot · 2026-10-08
- Frontier Models Decompiling Binaries Could Rescue Devices Bricked by Manufacturers — m4rkmc · 2026-10-08
- A different perspective: giving up on AI means aging will almost certainly kill us all — smith2008 · 2026-10-08
- State of AI 2026: frontier narrows to three labs, inference costs drop 13x a year — Nathan Benaich (Air Street) · 2026-10-08