Paper: How Much Do LLMs Remember? ~3.6 Bits Per Param
pmttyji · reddit · 2026-07-07
A new paper proposes a method to estimate a model's memory of individual data points to measure its capacity. The authors formally decompose memory into "incidental memory" (information specific to a dataset) and "generalization" (information about the true data generation process). By completely eliminating generalization, the total memory yields the capacity estimate: GPT-style models store approximately 3.6 bits per parameter.
When training on gradually increasing datasets, models first exhaust their capacity for memorization before entering the grokking phase, where incidental memory decreases and generalization increases. The researchers trained hundreds of transformers ranging from 500,000 to 1.5 billion parameters, deriving a series of scaling laws related to capacity, data scale, and membership inference.
Related event: ICML Paper: LLMs Store 3.6 Bits of Memory Per Parameter(2 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11