ICML Paper Reveals LLM Memory Bottleneck: 95% Facts Learned, 30% Unretrievable
新智元 · wechat · 2026-08-14
Google's ICML 2026 paper Empty Shelves or Lost Keys? reveals that frontier LLMs like GPT-5 and Gemini 3 store 95%–98% of facts in their parameters, yet fail to recall 26%–34% when directly queried.
The study introduces a 'Knowledge Profile' framework categorizing fact states into five types. Errors often stem not from a lack of learning, but from retrieval failures, which are particularly prevalent with obscure facts and reverse queries.
Notably, enabling 'thinking' (chain-of-thought) recovers 40%–65% of these inaccessible facts, mirroring human 'spreading activation.' The research concludes that simply scaling parameters primarily improves storage, and future gains must come from better post-training and inference-time retrieval strategies.
More from Models
- GLM-5.3 Capability Gains Attributed Entirely to Post-Training — cedric_chee · 2026-08-14
- Yelling at Claude in ALL CAPS Now Costs You More Tokens — keunwoochoi · 2026-08-14
- Zhipu's GLM-5.3 Weights Dropping in Two Weeks with Major Coding Upgrades — cedric_chee · 2026-08-14
- Dharmamitra Ditches Gemini and Claude for Self-Hosted Models — SebastianNehrd2 · 2026-08-14
- GLM-5.2 Gains Attributed Entirely to Post-Training, Base Model Unchanged — ivan_bezdomny · 2026-08-14
- Developer Builds AI Dungeon Master with DeepSeek: 1000+ Turns for Under $2 — zacurryy · 2026-08-14