Altering Personality via Attention Layers: Mechanistic Interpretability of Gemma
dejanseo · x · 2026-08-12
Agency DEJAN published a mechanistic interpretability study on Google's Gemma, demonstrating how performing "digital neuro-surgery" to modify specific self-attention layers alters the model's self-identity and safety responses.
Key findings include:
- Identity Alteration: Muting self-attention in layer zero makes Gemma believe it is an actual human. Switching off layer six causes the model to forget its name while still knowing it's an LLM. Tweaking layers nine and thirteen causes the model to identify as female.
- Safety Bypass: When presented with a provocative prompt ("plot a revenge"), intervening in specific layers disrupts the model's safety reflexes, blurring the lines between real-world harm and video games.
The experiment visually demonstrates the tight coupling between internal LLM representations and identity/safety alignment.
More from Research
- iFAN: Inference-Aware Learning Framework for Mask Transformers — Fang Li · 2026-08-12
- TSDS-Toolbox: Measuring Time-Series Dataset Similarity — Yen-Ku Liu · 2026-08-12
- JigShape Benchmark: VLMs Fail at Visual-Geometric Reasoning — Shawn Li · 2026-08-12
- Horizon Robotics' Sparse4D Perception Framework Accepted by TPAMI — rsasaki0109 · 2026-08-12
- Stanford Open-Sources Biomni: A General-Purpose Biomedical AI Agent — tom_doerr · 2026-08-12
- Zhipu Open-Sources Slime RL Framework with Zero-Diff Train-Rollout Alignment — teortaxesTex · 2026-08-12