800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining
The Pangram team released a scaling law study that pretrains roughly 800 language models to quantify the impact of the growing share of AI-generated text in web-crawled corpora on pretraining. Annotation with the Pangram detector found that in web data after FineWeb quality filtering, 27.5% of tokens were classified as AI-generated in June 2026, rising to 31.1% by August. For models with sufficient data, mixing in AI text actually increases the loss.
Confirmed
- The study pretrained 800 language models and examined the impact of AI-generated text directly on real web corpora, rather than only spiking in synthetic text.
- The share of AI-generated text keeps climbing: after FineWeb filtering, 27.5% in June and 31.1% in August (m3 mentioned an unfiltered figure of roughly 31%, consistent with the August data).
- Training on unfiltered crawled data requires about 3x the compute compared with training only on human text.
- The effect depends on data abundance: for data-scarce models with fewer than about 10 human tokens per parameter, adding AI-generated tokens lowers the loss on human text; but from the Chinchilla-optimal ratio (around 20) onward, adding more AI tokens raises the loss instead.
- Conclusion: for models already data-sufficient, training with AI-generated text mixed in is worse than adding no new data at all (as @MohitIyyer summarized in relaying @jennajrussell's thread).
Why it matters
- This study is the first to systematically quantify the "AI text backfire" effect at scale with real corpora and hundreds of models: as AI-generated content keeps flooding the internet, indiscriminately crawling it for training is not just useless — it significantly raises compute costs and degrades model quality.
- For small, data-scarce models, synthetic data still helps, showing the effect is conditional and providing a clear reference point for data-mixing strategies (the boundary between roughly 10 tokens/parameter and Chinchilla's 20).
2026-10-01 ~ 2026-10-02 · 5 related posts
Primary sources
- [source] Pretraining 800 LMs shows AI-generated web text can actively hurt scaling — iScienceLuvr · 2026-10-01
- New finding: AI-generated tokens raise loss once models pass Chinchilla-optimal data ratios — gleech · 2026-10-02
- New paper: training on AI-polluted web crawls costs up to 3x compute by 2028 — cephaloform · 2026-10-02
2 near-duplicate retellings: pangram · MohitIyyer