New paper: training on AI-polluted web crawls costs up to 3x compute by 2028

cephaloform · x · 2026-10-02

A new paper by Jenna Russell measures how AI-generated text in web crawls distorts LLM pretraining scaling laws. At August 2026's 31% AI share, training on the unfiltered crawl takes 1.6x the compute of training on its human-written portion; at a forecast 51% share by end of 2028, it will take 3.0x.

Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→

Original post →

More from Models

Models channel →