Overfiltering Wastes Compute; Data Selection Beats Simple Heuristics
giffmana · x · 2026-08-22
Arguing against naive data filtering, the author distinguishes between filtering and selection. Simple heuristics (e.g., filename rules) remove garbage but also discard valuable data (e.g., date-named images), wasting compute. In scenarios of limited resources or non-general models, data selection is preferred over crude filtering. The quoted researcher adds that the value lies more in importance weighting and selecting source data for rephrasing rather than precise filters.
Related event: Debate Over Data Filtering vs. Selection in Model Training(5 posts)→
More from Research
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24