COLM Paper: Meaningful Neuron Directions Extracted from MLP Weights Alone
megamor2 · x · 2026-10-08
At COLM 2025's poster session, Asaf Avrahamy, Yoav Gur and megamor2 presented their paper "Disentangling MLP Neuron Weights in Vocabulary Space".
Key result: the team shows that meaningful directions explaining how individual MLP neurons work can be extracted purely from weight data, without relying on activations — offering a weights-only route to mechanistic interpretability of vocabulary-space neuron roles.
More from Research
- SaTML 2027 to host AdvML Frontiers workshop on human-centered trustworthy machine learning — pinyuchenTW · 2026-10-08
- OpenAI's new result proves 2005 edit-distance embedding optimal; researcher distills proof to 2.5 pages with AI help — thegautamkamath · 2026-10-08
- Experiments show smarter models and higher effort write better LLM-judge evals — danshipper · 2026-10-08
- Tencent's WorkForge scales verifiable training environments for long-horizon work agents — teortaxesTex · 2026-10-08
- Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation — zhenjun_zhao · 2026-10-08
- DensiTok: flow-matching token densification lets frozen feed-forward 3DGS see unseen views — zhenjun_zhao · 2026-10-08