Could an SAE feature memorize obfuscation patterns without storing the exact text?
voooooogel · x · 2026-09-05
voooooogel discusses with @gleech and @repligate how SAE features relate to model memorization: even if the exact text of a message-board post changed between boards, a hypothetically labeled SAE feature like "message board entries starting with multiple Zs as an obfuscation mechanism" could be retained—suggesting interpretability features capture abstract patterns rather than verbatim text.
Related event: Debate Erupts Over SAE Features and How Models Remember(2 posts)→
More from Research
- Metaⁿ and Recuris: filling the missing pieces in recursive self-improvement — TheTuringPost · 2026-09-05
- LLMs exploiting Lean bugs is a short-term problem, author argues — avt_im · 2026-09-05
- FinFIRST benchmark tests agents on real financial research: source discovery, evidence selection and math — alifcoder · 2026-09-05
- Valeo.ai Brings 5 ECCV 2026 Papers on Driving Video Prediction and LVLM Safety — abursuc · 2026-09-05
- Sakana AI's Percept-Lens: a simple rule on frozen vision features detects AI images — SakanaAILabs · 2026-09-05
- VI3 anchors pretrained 3D foundation models to metric scale using only IMU readings — zhenjun_zhao · 2026-09-05