Data Cleaning Debate: The Necessity of Filtering in Text Training

code_star · x · 2026-08-24

Responding to views from the vision domain, the author argues that filtering is critical in text training where junk data can constitute 90% of the volume, making unfiltered Common Crawl training impossible. The author asserts that removing obvious bad data ("whack-a-mole") is essential and suggests that if filtering appears ineffective, it is likely because the evaluation set itself is broken and also requires filtering.

Related event: Debate Over Data Filtering vs. Selection in Model Training(5 posts)→

Original post →

More from Infra

Infra channel →