Simplex research derives transformers' "multidimensional features" activations from theory

livgorton · x · 2026-09-10

New work from Simplex asks what structure LLM activations should have, given training data drawn from many sources. Under natural assumptions about data structure, the belief geometry forms "telescoping cones," and transformers are shown to represent this structure. The post offers a theoretical derivation of the interpretability field's picture of activations as sparse linear combinations of multidimensional features, rather than relying purely on empirical observation.

Related event: Simplex Finds 'Telescoping Cone' Belief Geometry in LLM Activations(3 posts)→

Original post →

More from Research

Research channel →