COLM Paper: Meaningful Neuron Directions Extracted from MLP Weights Alone

megamor2 · x · 2026-10-08

At COLM 2025's poster session, Asaf Avrahamy, Yoav Gur and megamor2 presented their paper "Disentangling MLP Neuron Weights in Vocabulary Space".

Key result: the team shows that meaningful directions explaining how individual MLP neurons work can be extracted purely from weight data, without relying on activations — offering a weights-only route to mechanistic interpretability of vocabulary-space neuron roles.

Original post →

More from Research

Research channel →