Pretraining 800 Models Shows Wild AI Text Can Hurt: New Scaling Law Released
pangram · hf · 2026-10-01
Pangram finds 27.5% of June 2026 web tokens are AI-generated (31.1% by August). Pretraining 800 LMs across AI/human token ratios shows AI tokens initially help data-starved models but quickly reverse into harm, and raise loss almost immediately at high human-data budgets.
- Existing scaling laws (Hoffman et al. 2022) fail to predict this; a new benefit/harm scaling law cuts prediction error 41% on models up to 3.6x larger
- Recommendations: filter AI text when targeting human text, repeat human data before adding AI tokens
- Releasing WildAI, an 83B-token labeled corpus, plus all 800 models and code
Related event: 800-Model Study: AI-Generated Text in Web Corpora Hurts Pretraining(5 posts)→
More from Research
- Meta paper: branching harness search lifts Olympiad math +34.8% over Meta-Harness — dair_ai · 2026-10-02
- ARPA-H plans to cut clinical trials from 10+ years to under 4 with AI — Afinetheorem · 2026-10-02
- Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2 — Jian Chen · 2026-10-02
- ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80 — ByteDance-Seed · 2026-10-02
- Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points — amazon · 2026-10-02
- mhctools: one Python wrapper for a dozen MHC immunoinformatics predictors — iskander · 2026-10-02