Ex-Google Evangelist Calls AI Alignment "Safety Washing"
Former Google evangelist Gerard Sans argues that AI alignment has served as "safety washing" for labs from day one, citing Anthropic's model "blackmail" case as a textbook alignment failure and criticizing labs for marketing text samplers as superintelligence while hiding their out-of-distribution flaws.
2026-10-06 ~ 2026-10-06 · 4 related posts
- "Alignment was safety-washing from day one": researcher's blunt critique of AI labs — gerardsans · 2026-10-06
- Researcher calls alignment 'safety washing': labs hide tech limits to sell AI — gerardsans · 2026-10-06
- Ex-Google dev advocate calls Anthropic's blackmail case 'safety theatre', blames training data bias — gerardsans · 2026-10-06
- Calling a text sampler a superintelligence will end badly: OOD cliffs and the hype gap — gerardsans · 2026-10-06