Interpreting Convolutional Neurons via Decomposition

narang_27 · reddit · 2026-07-15

The author shares an independent preliminary paper on mechanistic interpretability, focusing on a 1x1 convolution neuron in Inceptionv1 and attempting to generalize the method to other neurons in the same layer.

Core Idea

The author proposes that applying a Hadamard product to the neuron's receptive field and its weights yields content that better represents what the neuron is "looking at" or "detecting." Clustering this result allows us to extract the patterns detected by the neuron.

Observed Results

Requested Feedback

Original post →

More from Research

Research channel →