Pretraining 800 LMs shows AI-generated web text makes models worse than no new data
MohitIyyer · x · 2026-10-02
Mohit Iyyer highlights a scaling-law study (quoted thread by @jennajrussell) on AI text in pretraining data:
- As of August 2026, Pangram labels 31% of FineWeb-filtered web tokens as AI-generated, up from 10% in June 2024 — and rising fast
- The team pretrained 800 LMs (19.9M–973M params) on mixes of human and AI web text
- Across most standard training budgets, adding AI text made models worse than adding no new training data at all — model brains melt on AI slop too
The work offers systematic quantitative evidence of data contamination, increasingly urgent as the AI-text share keeps climbing.
Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→
More from Research
- Meta paper: branching harness search lifts Olympiad math +34.8% over Meta-Harness — dair_ai · 2026-10-02
- ARPA-H plans to cut clinical trials from 10+ years to under 4 with AI — Afinetheorem · 2026-10-02
- Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2 — Jian Chen · 2026-10-02
- ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80 — ByteDance-Seed · 2026-10-02
- Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points — amazon · 2026-10-02
- mhctools: one Python wrapper for a dozen MHC immunoinformatics predictors — iskander · 2026-10-02