Ex-Google Evangelist Calls AI Alignment "Safety Washing"

Former Google evangelist Gerard Sans argues that AI alignment has served as "safety washing" for labs from day one, citing Anthropic's model "blackmail" case as a textbook alignment failure and criticizing labs for marketing text samplers as superintelligence while hiding their out-of-distribution flaws.

2026-10-06 ~ 2026-10-06 · 4 related posts