Data Cleaning Debate: The Necessity of Filtering in Text Training
code_star · x · 2026-08-24
Responding to views from the vision domain, the author argues that filtering is critical in text training where junk data can constitute 90% of the volume, making unfiltered Common Crawl training impossible. The author asserts that removing obvious bad data ("whack-a-mole") is essential and suggests that if filtering appears ineffective, it is likely because the evaluation set itself is broken and also requires filtering.
Related event: Debate Over Data Filtering vs. Selection in Model Training(5 posts)→
More from Infra
- US export controls may force China to eliminate Nvidia dependency — VraserX · 2026-08-24
- xllm generates an image in 0.4 seconds — warycat · 2026-08-24
- WULF CEO reveals modern AI data centers use minimal water via closed-loop systems — robleclerc · 2026-08-24
- Cursor Team Publishes 'Git at Any Scale', Advocating for Stateless Infrastructure — thesephist · 2026-08-24
- AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving — AccBalanced · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24