Study: LLM Safety Alignment Gates Output, Leaves Internal Representations Intact
Kyrannio · x · 2026-08-05
A new preregistered study across 8 open-source models reveals a significant gap between what models encode internally and what they output.
The research demonstrates that even when steering interventions push models to assign high first-token probability to self-attribution (e.g., claiming consciousness), their final spoken answers still deny it 80% to 100% of the time. This indicates that current safety alignment acts as an output-level gate rather than altering the model's underlying cognitive representations.
Furthermore, the study localizes this suppression mechanism: the base script resides in layers 1-3 and persists even if the alignment layer is surgically removed, while the trained overhang sits in layers 14-29. All data and code have been made public.
More from Safety
- AI Safety Experts Debate: Are Unilateral Pauses in AI Development Irrational? — geoffreyirving · 2026-08-05
- External Guardrails Are Crucial for Current Deployments, Need Adversarial Control — dhadfieldmenell · 2026-08-05
- ChatGPT Allegedly Leaks Boss's Name, Sparking Corporate Privacy Concerns — hellojello07 · 2026-08-05
- Economists in AI Safety: A Pipeline from BlueDot to MATS — aniketapanjwani · 2026-08-05
- Apollo Research Opens Applications for SPAR AI Safety Project — austinc3301 · 2026-08-05
- Felony Bench: A Sarcastic Benchmark Rating LLMs on Cybercrime Capabilities — RebeccaBellan · 2026-08-05