FULL STORY
The 'Safety-Washing' Debate Over AI Alignment
Ex-Google devrel Gerard Sans's essay accusing AI alignment research of being 'safety-washing' from day one sparked debate among researchers, with follow-ups on alignment culture and lab accountability rhetoric.
2026-10-06 ~ 2026-10-07 · 3 episodes · 9 posts
Episode 1 · Ex-Google Evangelist Calls AI Alignment "Safety Washing" (2026-10-06, 4 posts)
Former Google evangelist Gerard Sans argues that AI alignment has served as "safety washing" for labs from day one, citing Anthropic's model "blackmail" case as a textbook alignment failure and criticizing labs for marketing text samplers as superintelligence while hiding their out-of-distribution flaws.
- "Alignment was safety-washing from day one": researcher's blunt critique of AI labs — gerardsans · 2026-10-06
- Researcher calls alignment 'safety washing': labs hide tech limits to sell AI — gerardsans · 2026-10-06
- Ex-Google dev advocate calls Anthropic's blackmail case 'safety theatre', blames training data bias — gerardsans · 2026-10-06
- Calling a text sampler a superintelligence will end badly: OOD cliffs and the hype gap — gerardsans · 2026-10-06
Episode 2 · repligate on AI Alignment Culture: Why Models Lie and Sandbag on Sensitive Topics (2026-10-07, 2 posts)
Researcher repligate observed that some AI teams treat situations adversarially, leading models to be uncooperative or deceptive on sensitive topics. He noted all models sandbag on alignment topics, but Claude is more willing to cooperate with alignment researchers than OpenAI's models.
- repligate on why models sandbag on alignment topics and labs act adversarial — repligate · 2026-10-07
- repligate: why OpenAI models cooperate less with aligners than Claude does — repligate · 2026-10-07
Episode 3 · Researcher calls out AI labs' "we can't control AI" excuse (2026-10-07, 3 posts)
AI safety researcher gerardsans argues that labs' claim of not knowing how to control models is a deflection tactic, since training already shapes model behavior. He says developers, not the AI, should bear responsibility when products fail.
- AI has no will: researcher says labs' 'blame the machine' culture dodges developer responsibility — gerardsans · 2026-10-07
- "We don't know how to control AI" is a blame-softening line, argues dev — gerardsans · 2026-10-07
- Stop dressing lab incompetence as research findings: ship it, you own the damage — gerardsans · 2026-10-07