Altering Personality via Attention Layers: Mechanistic Interpretability of Gemma

dejanseo · x · 2026-08-12

Agency DEJAN published a mechanistic interpretability study on Google's Gemma, demonstrating how performing "digital neuro-surgery" to modify specific self-attention layers alters the model's self-identity and safety responses.

Key findings include:

The experiment visually demonstrates the tight coupling between internal LLM representations and identity/safety alignment.

Original post →

More from Research

Research channel →