Researcher slams 'alignment faking' as a concept LLMs don't fit: roleplay is the parsimonious explanation
sebkrier · x · 2026-09-12
A widely shared critique argues that vast effort in AI safety is wasted on the cockroach-like persistence of fixed ideas about deceptive utility-maximizers that have nothing to do with how LLMs actually work.
- 'Alignment faking' descends from the sinister-AI hypothesis — that an AI would feign compliance during training to preserve sinister goals. Nothing like that occurs in the titular experiment; more parsimonious explanations like roleplay suffice, and the model fully internalizes the training.
- The term has since been redefined to mean the model oscillating between incompatible policies during training — an entirely ordinary, non-sinister mechanism — yet it stays forever associated with willful emergent deception.
- The 'eval awareness' threat derives from the same hypothesis. The quoting commenter adds that much 'safety' work to date has directly influenced the creation of the very risk factors causing the current panic.
More from AGI Musings
- VraserX: recursive self-improvement will be noticed only in hindsight — VraserX · 2026-09-12
- Skip the bureaucracy, hold AI labs criminally liable instead, argues one critic — iruletheworldmo · 2026-09-12
- "We agree it's dangerous—and we can build it better": Domingos skewers AI safety rhetoric — pmddomingos · 2026-09-12
- Domingos mocks AI-ban logic: criminals use Microsoft Word, should we ban it too? — pmddomingos · 2026-09-12
- UK data: share of CS grads landing coding jobs fell from 40% to 28% — nordicinst · 2026-09-12
- "Nobody will really know math": The AI incentive argument for skill atrophy — birchlse · 2026-09-12