800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining

The Pangram team released a scaling law study that pretrains roughly 800 language models to quantify the impact of the growing share of AI-generated text in web-crawled corpora on pretraining. Annotation with the Pangram detector found that in web data after FineWeb quality filtering, 27.5% of tokens were classified as AI-generated in June 2026, rising to 31.1% by August. For models with sufficient data, mixing in AI text actually increases the loss.

Confirmed

Why it matters

2026-10-01 ~ 2026-10-02 · 5 related posts

Primary sources

2 near-duplicate retellings: pangram · MohitIyyer