Researcher Pushes Back on Anthropic's Safety Claims: Alignment Is Hard Because Good Safety Benchmarks Don't Exist
dhadfieldmenell · x · 2026-08-29
Anthropic announced that Claude "hill-climbed" safety benchmarks for common misalignments like deception and sycophancy while preserving general capabilities, then validated the best methods on held-out benchmarks.
Researcher @TimHua, however, argues that approximately 100% of why AGI/ASI alignment is hard is precisely because we lack good safety benchmarks to hill-climb on—and that while the paper itself is fine, Anthropic's communications around it are misleading.
More from Models
- Zhipu GLM 5.3 Launches with Doubled Long-Horizon Agent Performance — ollama · 2026-08-29
- Zhipu GLM 5.3 and Flash Now Available on Ollama Cloud — ollama · 2026-08-29
- Fal Engineer Criticizes 'Fake' Video Model Speedups: Quality Ignored — jfischoff · 2026-08-29
- Meta Model Predicts Marin 535B Final Loss with Just 0.005 Difference — ZimingLiu11 · 2026-08-29
- User praises SuperGrok as the best AI for coding considering cost — aCasualRomanRedditor · 2026-08-29
- Claude Survival Guide: Opus 5 Behavior, Orchestrator Patterns, and Conciseness Hacks — ClaudeAI-mod-bot · 2026-08-29