New Paper Asks Whether SAEs Capture Concept Manifolds, Finds a 'Dilution' Failure Mode
aryaman2020 · x · 2026-09-05
In a discussion with bearseascape about what the right level of abstraction for circuits is, aryaman2020 cites an arXiv paper by Usha Bhalla et al., 'Do Sparse Autoencoders Capture Concept Manifolds?'
- The question: SAEs implicitly assume concepts are independent linear directions, but growing evidence suggests many concepts are organized along low-dimensional manifolds encoding continuous geometric relationships.
- Framework: the paper defines what it means for an SAE to capture a manifold and shows two modes — global (a compact group of atoms whose linear span contains the manifold) and local (features that each tile a restricted region of the geometry).
- Empirical result: SAEs recover continuous structures suboptimally, mixing both solutions in a fragmented regime the authors call 'dilution,' explaining why manifold structure is rarely visible at the level of individual concepts.
- Discussion: aryaman2020 argues that if using neurons for circuits, you'd ideally group granular neurons into sets writing to complementary output classes, but superposition may make things too messy — and points to an Ising-model-style framing.
He agrees it's unclear what the best level of abstraction for a circuit is.
Related event: Circuit interpretability hits a wall as ablations miss across tasks(4 posts)→
More from Research
- IndianRailwayBench ranks LLMs by their ability to book tatkal train tickets — Paimaamu · 2026-09-06
- NEAR AI's open-source Lean agent solves all of Putnam Bench for just $111 — lukaszkaiser · 2026-09-06
- Russian startup Mostik bridges LLM hidden states, cutting cost to 1/20 — 机器之心 · 2026-09-06
- KV Cache Explained: Why It's Crucial in LLM Inference and Often Misunderstood — techNmak · 2026-09-06
- PhD Student Uses Multi-Agent AI to Crack a 98-Year-Old Math Problem in 48 Hours — 量子位 · 2026-09-06
- New piece: Cognitive maps as a medium for thought — abenitezburraco · 2026-09-06