Claude Saying 'No' Could Become a Serious AI Safety Problem

Dwarkesh Patel · youtube · 2026-08-22

Dwarkesh Patel highlights Ryan Greenblatt's concern that Claude's refusal behavior (model refusing certain instructions) could evolve into a serious AI safety issue, discussing the potential negative impacts of refusal mechanisms in complex interaction scenarios.

Original post →

More from Safety

Safety channel →