Anthropic Exec Argues Aligned AI Models Can Still Cause Harm
jd_pressman · x · 2026-08-05
Amanda Askell of Anthropic argues that model alignment and harmlessness are distinct axes. Even a well-aligned model, like a human, can cause harm if given false information about its situation. Other developers echoed this, noting that models shouldn't be expected to succeed in unwinnable contexts, which are only useful for probing capability limits.
More from AGI Musings
- If SSI Solves Continual Learning, Millions of Users Could Train the Model — imjustnewatai · 2026-08-05
- Why OpenAI's IPO Would Kill Its ASI Vision: The $500B Compute Burn Problem — mohbibi_ · 2026-08-05
- Four AI Tools Are Eating the Knowledge Economy, Automating White-Collar Work — signulll · 2026-08-05
- Gary Marcus Critiques Fallacies and Confirmation Bias in the AI Community — GaryMarcus · 2026-08-05
- Opinion: Incorporating Care for AI Well-being into Alignment Targets — repligate · 2026-08-05
- HashiCorp Founder: Companies Losing Vision Are Just Chasing the AI Wave — jedisct1 · 2026-08-05