Mechanistic Interpretability Test: Altering Gemma's Preferences via Specific MLP Layers
dejanseo · x · 2026-08-04
After upgrading his mechanistic interpretability stack for Gemma models, the author shared a recent finding. By applying a temporary hook at scale 0 on model.layers.12.mlp and model.layers.13.mlp, and reading a specific token (unembed.tok#9595), he successfully altered the model's inherent preferences (e.g., its "favorite color is no longer blue"). He noted this is commercial research and withheld the strategic intent but promised to share more findings soon.
More from Research
- Nature Study: AI Dermatology Diagnosis Amplifies Public Automation Bias — EricTopol · 2026-08-04
- Radical Co-founder on Training the Largest Genome Model to Write DNA — exnx · 2026-08-04
- 3 Lines of Code Fixed 123 Failed PPO Experiments by Changing Reward Shaping — mikeysce · 2026-08-04
- Research: Frozen Pixel-Space Diffusion Models Can Self-Guide — nanyang-technological-university-singapore · 2026-08-04
- ICDAR 2026 Competition: Multimodal AI Struggles with Scientific Figures — SciKnowOrg · 2026-08-04
- GEOID-Flood: Large-Scale Multi-Modal Benchmark for Flood Segmentation — links-ads · 2026-08-04