New paper: training on AI-polluted web crawls costs up to 3x compute by 2028
cephaloform · x · 2026-10-02
A new paper by Jenna Russell measures how AI-generated text in web crawls distorts LLM pretraining scaling laws. At August 2026's 31% AI share, training on the unfiltered crawl takes 1.6x the compute of training on its human-written portion; at a forecast 51% share by end of 2028, it will take 3.0x.
Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→
More from Models
- JevBench adds multilingual queries to test Jev models across languages — airesearch12 · 2026-10-02
- Perplexity open-sources multimodal decision model at $0.04 per million input tokens — AravSrinivas · 2026-10-02
- Grok 4.7 rolls out in the Grok app for chat and research after long wait — mark_k · 2026-10-02
- OpenAI's GPT-6 Astra is shockingly good at controlling robots — binarybits · 2026-10-02
- Reddit user claims wity-1 decision model beats Jev on all 4 benchmarks, tops image bench — boneMechBoy69420 · 2026-10-02
- Dev says Sol 6.1 is painfully slow compared to Opus 5.5 — weswinder · 2026-10-02