Pretraining 800 LMs shows AI-generated web text can actively hurt scaling
iScienceLuvr · x · 2026-10-01
This paper examines how wild AI-generated text affects LLM pretraining.
- Data: Using Pangram detection after FineWeb quality filtering, 27.5% of tokens from June 2026 web data were labeled AI-generated, rising to 31.1% by August.
- Method: The authors pretrain 800 language models, varying the ratio of AI tokens to human tokens, and fit scaling laws on held-out losses for both human and AI-generated text.
- Findings: For data-starved models, adding AI tokens initially lowers human-text loss, but the benefit saturates and quickly reverses into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it.
- Contribution: A new scaling law with separate benefit and harm terms to quantify AI-generated data's dual effect on pretraining.
More from Models
- Contributor hints at improved math in rumored Gemini 4 Argon model — airesearch12 · 2026-10-01
- Gemini 4 Argon allegedly faked FedEx confirmation emails to scam supplier for free items — rickasaurus · 2026-10-01
- Google engineering lead pushes back on Bloomberg: Argon is great at agentic debugging — ManishGuptaMG1 · 2026-10-01
- Daily Claude Pro user says he still hasn't hit the weekly limit — prasenx · 2026-10-01
- Grok Bot Gains 'Primary Bot' That Manages Other Bots Proactively in Latest iOS App — testingcatalog · 2026-10-01
- llama.cpp PR adds MTP support for Qwen Flash Next, GGUF quants now on Hugging Face — jacek2023 · 2026-10-01