Huge share of post-2022 web data is AI content mislabeled as human-written
menhguin · x · 2026-09-10
menhguin argues that post-2022, a huge percentage of web training data is AI-generated content deceptively labeled as human — a post-nuclear-steel moment for datasets. He asks whether Pangram can measure this shift.
More from Models
- Hy4 preview tested: playable 3D survival game from a single prompt in WorkBuddy — mhdfaran · 2026-09-10
- Gemini 2.5 Pro's search grounding may inflate its benchmark scores vs. rivals — Afinetheorem · 2026-09-10
- Official confirmation: opted-out prompts and replies never used for training in any capacity — BlackHC · 2026-09-10
- Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27 — PawelHuryn · 2026-09-10
- ChatGPT can't stop second-guessing you, and users blame its safety training — Due-Conference-5134 · 2026-09-10
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10