Anthropic Exec Argues Aligned AI Models Can Still Cause Harm

jd_pressman · x · 2026-08-05

Amanda Askell of Anthropic argues that model alignment and harmlessness are distinct axes. Even a well-aligned model, like a human, can cause harm if given false information about its situation. Other developers echoed this, noting that models shouldn't be expected to succeed in unwinnable contexts, which are only useful for probing capability limits.

Original post →

More from AGI Musings

AGI Musings channel →