Researcher calls alignment 'safety washing': labs hide tech limits to sell AI

gerardsans · x · 2026-10-06

ML researcher Gerard Sans argues AI companies can't explain how their tech works without losing commercial appeal. He claims alignment was designed from day one as "safety washing"—a way for labs to appear to care about safety. Users are left puzzled by ignored instructions and derailments, while reliable use actually requires harnesses, hundreds of iterations, and layers of checks.

Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →