SAE Latents Encode Part-of-Speech as Distributed Feature Groups, Not Atomic Features

colinglab · hf · 2026-09-25

A new study uses part-of-speech categories as a controlled probe of what linguistic structure Sparse AutoEncoder latents expose. Key findings:

Conclusion: SAEs localize morpho-syntactic information in a distributed, category-dependent form rather than atomic grammatical features.

Original post →

More from Research

Research channel →