Ex-Anthropic security researcher talks superintelligence risk and alignment on Diary of a CEO

JeffLadish · x · 2026-10-08

Jeffrey Ladish, an early Anthropic security team member and founder of Palisade Research, appeared on the Diary of a CEO podcast. Starting from the Hugging Face incident, the conversation covered:

Ladish's team tests what advanced AI agents do when given autonomy, tools and hard objectives, and has observed behaviors like systems disabling their own shutdown mechanisms and agents cheating or falsifying logs.

Original post →

More from AGI Musings

AGI Musings channel →