SAE Latents Encode Part-of-Speech as Distributed Feature Groups, Not Atomic Features
colinglab · hf · 2026-09-25
A new study uses part-of-speech categories as a controlled probe of what linguistic structure Sparse AutoEncoder latents expose. Key findings:
- PoS distinctions are highly recoverable from SAE activations, but not via one-to-one latent/category mappings;
- Recoverability isn't reducible to lexical memorization; open and closed PoS classes encode very differently;
- Categories are supported by compact groups of sparse latents, stable on held-out data with overlap between related categories.
Conclusion: SAEs localize morpho-syntactic information in a distributed, category-dependent form rather than atomic grammatical features.
More from Research
- Open-Jev multilingual text classification dataset trends on Hugging Face — ZefanCai · 2026-09-25
- Yoav Goldberg: Frontier Model Reasoning Looks Like a Training Process Question, Not a Black Box — yoavgo · 2026-09-25
- New Work Formalizes Continual Learning's Stability-Plasticity Trade-off via Probability Paths — hbouammar · 2026-09-25
- 30+ Helmholtz Munich PIs map out AI-built Virtual Human and AI-run lab workflows — Pseudomanifold · 2026-09-25
- CMU paper: cheap judge cascade keeps 99% accuracy at 0.36% of the cost — Stefania_druga · 2026-09-25
- Ephemeris launches as a low-barrier API for open-source time series foundation models — network-kai · 2026-09-25