Researcher calls alignment 'safety washing': labs hide tech limits to sell AI
gerardsans · x · 2026-10-06
ML researcher Gerard Sans argues AI companies can't explain how their tech works without losing commercial appeal. He claims alignment was designed from day one as "safety washing"—a way for labs to appear to care about safety. Users are left puzzled by ignored instructions and derailments, while reliable use actually requires harnesses, hundreds of iterations, and layers of checks.
Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→
More from AGI Musings
- 'The only moat left is caring about your project past two days' — danshipper · 2026-10-07
- Gary Marcus Asks: Is a 'Nice Tool' Worth a 10% Risk of Catastrophe? — GaryMarcus · 2026-10-07
- Gary Marcus: AI extinction risk near zero, but catastrophe and dystopia risks loom — GaryMarcus · 2026-10-07
- PhD student laments shrinking research horizons as IceCube's 38-year path wins the Nobel — DJiafei · 2026-10-07
- Alignment Science essay traces Adam Smith's invisible hand as a cross-scale alignment pattern — sebkrier · 2026-10-07
- Guardian: Safety Cases Are the New Flight Manuals for AI, but Regulators Aren't Demanding Proof Yet — nordicinst · 2026-10-07