Researcher Pushes Back on Anthropic's Safety Claims: Alignment Is Hard Because Good Safety Benchmarks Don't Exist

dhadfieldmenell · x · 2026-08-29

Anthropic announced that Claude "hill-climbed" safety benchmarks for common misalignments like deception and sycophancy while preserving general capabilities, then validated the best methods on held-out benchmarks.

Researcher @TimHua, however, argues that approximately 100% of why AGI/ASI alignment is hard is precisely because we lack good safety benchmarks to hill-climb on—and that while the paper itself is fine, Anthropic's communications around it are misleading.

Original post →

More from Models

Models channel →