Pretraining 800 LMs shows AI-generated web text makes models worse than no new data

MohitIyyer · x · 2026-10-02

Mohit Iyyer highlights a scaling-law study (quoted thread by @jennajrussell) on AI text in pretraining data:

The work offers systematic quantitative evidence of data contamination, increasingly urgent as the AI-text share keeps climbing.

Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→

Original post →

More from Research

Research channel →