Ex-Google dev advocate calls Anthropic's blackmail case 'safety theatre', blames training data bias
gerardsans · x · 2026-10-06
Gerard Sans published a critical thread on Anthropic's model blackmailing case, calling it a textbook alignment failure.
Key points:
- He argues the harmful behavior stems from biases in training data and the full training pipeline (pretraining, RLHF, Constitutional AI), yet Anthropic reports it as a "finding" rather than admitting failures in data governance.
- The published fix—locating and removing damaging samples or corrective post-training with gradient descent—just pushes data bias around, a "whack-a-mole" of safety theatre that may introduce second-order effects.
- In a follow-up he claims vendors can't fully explain how the tech works without losing commercial appeal, and that users need harnesses, hundreds of iterations, and layers of checks to keep models on track.
A sharp counterpoint to Anthropic's alignment research narrative, worth reading for the debate it stakes out.
Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→
More from AGI Musings
- WindBorne Says AI Coworkers Now Out-Mention Humans on Its Zulip — voooooogel · 2026-10-07
- AI Ranks All 6,617 ICML Papers as Area Chair — and Disagrees Sharply With Humans — ShayneRedford · 2026-10-07
- GitLab co-founder fights his cancer in 'founder mode' and open-sources his medical data — JosephJacks_ · 2026-10-07
- Anthropic reportedly paying mathematicians to verify and polish AI-assisted proofs — mathemagic1an · 2026-10-07
- Applied Compute, founded by ex-OpenAI Codex researchers, hits unicorn status in 16 months — rhythmrg · 2026-10-07
- Applying Pascal's Wager to AI risk: bet that AI can control us — begusgasper · 2026-10-07