Most recurring multilingual SAE features fail causal validation — only one truly drives translation in Gemma
ravithejads · x · 2026-09-08
A BlackboxNLP 2026 Special Track paper causally validates multilingual sparse autoencoder (SAE) translation features in Gemma 2 and Gemma 3:
- Reproduces Wu et al.'s SAE feature discovery and extends it across varying prompt/source/target languages
- In both models, 20+ features fire frequently across all discovery settings, but causal validation (amplifying/ablating activations) shows nearly all have small or inconsistent effects — feature recurrence can overstate cross-lingual transfer
- The exception: Gemma 2's (L10, 5717) and Gemma 3's (L20, 2456) consistently improve COMET scores when amplified and degrade when ablated across 23 language settings — a language-agnostic translation-initiation direction
A methodological warning for interpretability researchers: activation frequency alone is unreliable; causal validation is a must.
More from Research
- What 'blowup' actually means: your coffee cup is a $1M math problem — burny_tech · 2026-09-08
- Armory serves VLA policies to 14 robots from one cloud GPU at 30Hz, accepted at CoRL 2026 — danfei_xu · 2026-09-08
- ECCV 2026 talk to dissect representations inside multimodal foundation models — abursuc · 2026-09-08
- Mathematician: Astra solved in 70 minutes an asymptotic that stumped us for a week — arampell · 2026-09-08
- Turn PRs into RL environments at scale — startups already raised seed rounds on this repo — lvwerra · 2026-09-08
- Paper finds CoT reasoning operations are geometrically organized in hidden states — dair_ai · 2026-09-08