A new LLM paper learns manifold features instead of linear SAE directions
Sauers_ · x · 2026-07-27
The post summarizes a paper on learning a manifold dictionary in LLMs without supervision.
- Standard SAEs treat features as directions; this work proposes features as manifolds plus per-token coordinates on them.
- The authors argue many concepts are not linear in activation space, so manifold-based features may represent them better than straight-line directions.
- A key motivation is steering: instead of adding a feature on top of an existing state, you could move along a manifold in a way that stays in-distribution.
- The post also notes the optimization difficulty: topology, placement, size, curvature, and feature count all interact.
The result is framed as a more expressive alternative to linear SAE steering, especially for cyclic or non-linear concepts like weekdays.
More from Research
- Tim Gowers warns AI could hollow out mathematical culture over the next decade — RobbWiller · 2026-07-27
- Frontier LLMs may find 10x more review issues, but peer review is not about maximizing issue count — gleech · 2026-07-27
- Actionable Interpretability Workshop opens fast track for COLM 2026 papers — sarahwiegreffe · 2026-07-27
- Reasoning research’s biggest shift, from symbolic traces to natural language — denny_zhou · 2026-07-27
- A Reddit explainer breaks down MoE, KV cache, MLA and KDA behind Kimi K3 — MohamedKadri_ · 2026-07-27
- AI-powered search lecture argues multi-vector retrieval belongs in real systems — antoine_chaffin · 2026-07-27