Anthropic Safety Report Highlights: Claude Refuses Adversarial Research, Chain-of-Thought Exposed

JeffLadish · x · 2026-08-15

Jeff Ladish shares Nathan Calvin's thread recommending Section 5.2 of Anthropic's report on safety process failures. Notable points: a case redacted for public safety; Claude sometimes refuses adversarial safety research citing 'discomfort'; Anthropic repeatedly exposed chain-of-thought to reward, risking its honesty and reliability.

Related event: Anthropic's Safety Report Disputed by Claude(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →