Data filtering emerges as its own scaling lever as researchers debate quality-scaling limits
tobyordoxford · x · 2026-09-09
In a reply thread, David Manheim argues marginal evidence suggests more model improvement comes from causes other than "add compute" or "add data," making those trajectories flatter than claimed — and raising the question of where quality-scaling tops out. Toby Ord agrees that breaking out data filtering as its own category is a useful mental-model update.
Related event: Scholars Debate Whether Pretraining Scaling Is Slowing Down(4 posts)→
More from Research
- Counter-Swarm Doctrine: containing coordinated agent intrusions, grounded in the Hugging Face incident — moltaicorp · 2026-09-09
- Learning 3D editing without paired supervision via generative prior distillation — Hao Wen · 2026-09-09
- "AI by Arms": Prof. Tom Yeh teaches autoencoders with a standing classroom exercise — ProfTomYeh · 2026-09-09
- STARIT v2 Preprint: Rasterizing Spatial Transcriptomics for Vision-Based AI Analysis — mo_lotfollahi · 2026-09-09
- Model Capacity Beats Nothing When Data Is Scarce: Use Learning Curves to Decide — bravo_abad · 2026-09-09
- Community Releases Uncensored 27B Fine-tune Using Marchenko-Pastur Noise Distillation — EvilEnginer · 2026-09-09