Mechanistic Interpretability Test: Altering Gemma's Preferences via Specific MLP Layers

dejanseo · x · 2026-08-04

After upgrading his mechanistic interpretability stack for Gemma models, the author shared a recent finding. By applying a temporary hook at scale 0 on model.layers.12.mlp and model.layers.13.mlp, and reading a specific token (unembed.tok#9595), he successfully altered the model's inherent preferences (e.g., its "favorite color is no longer blue"). He noted this is commercial research and withheld the strategic intent but promised to share more findings soon.

Related event: Upgraded Mechanistic Interpretability Allows Preference Modification in Gemma(2 posts)→

Original post →

More from Research

Research channel →