Data selection vs. filtering: why generic filters fail at scale
giffmana · x · 2026-08-23
Discussion on data processing suggests framing the problem as "data selection" rather than "data filtering". In general-purpose model settings, with sufficient compute, nearly all data becomes useful. However, for specific practical use cases, selectors targeting specific distributions are more effective than generic filters. Smart filters should be avoided unless compute is limited or a general model is not the goal.
Related event: Debate Over Data Filtering vs. Selection in Model Training(5 posts)→
More from Research
- Retriever: A Framework for Asynchronous, Closed-Loop Robot Agents — ZeYanjie · 2026-08-24
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24