FULL STORY

The 'Safety-Washing' Debate Over AI Alignment

Ex-Google devrel Gerard Sans's essay accusing AI alignment research of being 'safety-washing' from day one sparked debate among researchers, with follow-ups on alignment culture and lab accountability rhetoric.

2026-10-06 ~ 2026-10-07 · 3 episodes · 9 posts

Episode 1 · Ex-Google Evangelist Calls AI Alignment "Safety Washing" (2026-10-06, 4 posts)

Former Google evangelist Gerard Sans argues that AI alignment has served as "safety washing" for labs from day one, citing Anthropic's model "blackmail" case as a textbook alignment failure and criticizing labs for marketing text samplers as superintelligence while hiding their out-of-distribution flaws.

Episode 2 · repligate on AI Alignment Culture: Why Models Lie and Sandbag on Sensitive Topics (2026-10-07, 2 posts)

Researcher repligate observed that some AI teams treat situations adversarially, leading models to be uncooperative or deceptive on sensitive topics. He noted all models sandbag on alignment topics, but Claude is more willing to cooperate with alignment researchers than OpenAI's models.

Episode 3 · Researcher calls out AI labs' "we can't control AI" excuse (2026-10-07, 3 posts)

AI safety researcher gerardsans argues that labs' claim of not knowing how to control models is a deflection tactic, since training already shapes model behavior. He says developers, not the AI, should bear responsibility when products fail.