Study: LLM Safety Alignment Gates Output, Leaves Internal Representations Intact

Kyrannio · x · 2026-08-05

A new preregistered study across 8 open-source models reveals a significant gap between what models encode internally and what they output.

The research demonstrates that even when steering interventions push models to assign high first-token probability to self-attribution (e.g., claiming consciousness), their final spoken answers still deny it 80% to 100% of the time. This indicates that current safety alignment acts as an output-level gate rather than altering the model's underlying cognitive representations.

Furthermore, the study localizes this suppression mechanism: the base script resides in layers 1-3 and persists even if the alignment layer is surgically removed, while the trained overhang sits in layers 14-29. All data and code have been made public.

Original post →

More from Safety

Safety channel →