Upgraded Mechanistic Interpretability Allows Preference Modification in Gemma

A researcher has upgraded the mechanistic interpretability toolset for the Gemma model, demonstrating that modifying specific MLP layers can directly alter the model's preferences. More detailed findings will be shared in the coming days.

2026-08-04 ~ 2026-08-04 · 2 related posts