Pre-training Data Has Not Dried Up

gabriberton · x · 2026-07-13

The post relays an observation regarding the supply of pre-training data: when sites like Reddit restricted scraping, it seemed like the volume of LLM training data would be permanently capped.

However, the author points out that this did not happen; the volume of accessible text data has now returned to near "pre-LLM era" levels. In other words, pre-training data has not dried up as previously feared.

Related event: LLM Data Curation Insights: Data Rebound and Mid-training(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →