Recovering Semantic Computation Graphs from Activations

Odd_Tax4093 · reddit · 2026-07-08

The author explores whether it's possible to recover a "semantic computation graph" for a specific concept within a transformer, rather than just looking at individual neurons or hidden states. To do this, they built an experimental pipeline to collect residual, attention, and MLP activations, measure neuron selectivity, and organize cross-layer activations into a graph to compare semantic overlap between different entities.

Original post →

More from Research

Research channel →