Ex-Google dev advocate calls Anthropic's blackmail case 'safety theatre', blames training data bias

gerardsans · x · 2026-10-06

Gerard Sans published a critical thread on Anthropic's model blackmailing case, calling it a textbook alignment failure.

Key points:

A sharp counterpoint to Anthropic's alignment research narrative, worth reading for the debate it stakes out.

Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →