Modern LMs train on 100x more text than parameter capacity, debunking memorization myths

stanfordnlp · x · 2026-08-27

Refuting the idea that language models work by rote memorization, the post notes that modern training regimes ingest two orders of magnitude more text than can be stored in their parameters, making verbatim memorization physically impossible. It also points out that stored information is relatively inefficient. A free standard textbook on how LMs actually work is recommended.

Original post →

More from Research

Research channel →