Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization
voooooogel · x · 2026-09-05
A technical debate on LLM memorization and SAE interpretability. The author is skeptical that models verbatim-memorize training data in ways measurable by string edit distance — what he's seen is more like fuzzy recall of concepts.
Example: a model likely won't remember the exact text of a specific message-board post, but could hypothetically carry an SAE feature like 'writing message-board entries starting with multiple Z as an obfuscation mechanism' — abstract patterns rather than exact strings. The thread also touches on model size and scaling behavior in RL.
Related event: Debate Erupts Over SAE Features and How Models Remember(2 posts)→
More from Models
- Astra burns tokens at high effort for no gains, finds dev testing Terminal Bench 4.0 — zainhas · 2026-09-05
- GPT 6 Astra lets you change reasoning effort mid-conversation without breaking the cache — intellectronica · 2026-09-05
- "Don't use past tense for models": users mourn Claude Opus 3 — repligate · 2026-09-05
- Astra hits 74% on DeepSWE with 30k tokens, half the cost steps of GPT-5.6 Sol — haider1 · 2026-09-05
- Astra's C++ code called 'nonhuman': unreadable density, not superhuman — teortaxesTex · 2026-09-05
- Prompting Fable 5.1 with embodied mannerisms like *shrugs* makes roleplay flow better — repligate · 2026-09-05