Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization

voooooogel · x · 2026-09-05

A technical debate on LLM memorization and SAE interpretability. The author is skeptical that models verbatim-memorize training data in ways measurable by string edit distance — what he's seen is more like fuzzy recall of concepts.

Example: a model likely won't remember the exact text of a specific message-board post, but could hypothetically carry an SAE feature like 'writing message-board entries starting with multiple Z as an obfuscation mechanism' — abstract patterns rather than exact strings. The thread also touches on model size and scaling behavior in RL.

Related event: Debate Erupts Over SAE Features and How Models Remember(2 posts)→

Original post →

More from Models

Models channel →