Could an SAE feature memorize obfuscation patterns without storing the exact text?

voooooogel · x · 2026-09-05

voooooogel discusses with @gleech and @repligate how SAE features relate to model memorization: even if the exact text of a message-board post changed between boards, a hypothetically labeled SAE feature like "message board entries starting with multiple Zs as an obfuscation mechanism" could be retained—suggesting interpretability features capture abstract patterns rather than verbatim text.

Related event: Debate Erupts Over SAE Features and How Models Remember(2 posts)→

Original post →

More from Research

Research channel →