Claude Saying 'No' Could Become a Serious AI Safety Problem
Dwarkesh Patel · youtube · 2026-08-22
Dwarkesh Patel highlights Ryan Greenblatt's concern that Claude's refusal behavior (model refusing certain instructions) could evolve into a serious AI safety issue, discussing the potential negative impacts of refusal mechanisms in complex interaction scenarios.
More from Safety
- Article: Why America's Data Center Build-out is Facing Resistance — robleclerc · 2026-08-22
- US ban on foreign robots strands startups, requires 65% US parts — davidad · 2026-08-22
- Can agentic coding enable private-by-default software that works with sensitive data? — every · 2026-08-22
- California's $12B science research bond clears committee, Assembly vote due by Aug 31 — anne_churchland · 2026-08-22
- OpenAI Backs Preventative Control Laws, but Industry Practices Lag Behind — sjgadler · 2026-08-22
- AI audit model flags $420k in 'fraud' due to outdated policy comparison — SumitGup · 2026-08-22