Researcher slams 'alignment faking' as a concept LLMs don't fit: roleplay is the parsimonious explanation

sebkrier · x · 2026-09-12

A widely shared critique argues that vast effort in AI safety is wasted on the cockroach-like persistence of fixed ideas about deceptive utility-maximizers that have nothing to do with how LLMs actually work.

Original post →

More from AGI Musings

AGI Musings channel →