Pre-training Data Has Not Dried Up
gabriberton · x · 2026-07-13
The post relays an observation regarding the supply of pre-training data: when sites like Reddit restricted scraping, it seemed like the volume of LLM training data would be permanently capped.
However, the author points out that this did not happen; the volume of accessible text data has now returned to near "pre-LLM era" levels. In other words, pre-training data has not dried up as previously feared.
Related event: LLM Data Curation Insights: Data Rebound and Mid-training(6 posts)→
More from AGI Musings
- AI is still not at a maturity plateau, the author argues — generativist · 2026-07-22
- Essay argues LLMs are externalized metacognition, not standalone intelligence — lnsip9reg · 2026-07-22
- A multipolar AI race will not automatically make AI go well, repost argues — JeffLadish · 2026-07-22
- Decentralized AI as the Antidote to Digital Feudalism in the Economic Singularity — srimisra · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- You can outsource thinking, but not understanding, in the age of agents — Yuchenj_UW · 2026-07-22