AutoData: EMNLP paper uses agentic search for pre-training data selection
PMinervini · x · 2026-09-18
A new EMNLP paper, AutoData, explores whether autoresearch agents can find better pre-training data, not just better training code. The authors note data engineering gets far less attention than algorithmic autoresearch. The approach is simple—write heuristics to select data based on pre-constructed features—but works surprisingly well.
Related event: AutoData: Agentic Search for Better Pre-training Data Selection(2 posts)→
More from Research
- Berkeley's PixelRAG hits 10k stars: screenshots beat text parsing for RAG retrieval — tom_doerr · 2026-09-19
- Terence Tao's SAIR nonprofit unveils Open Math Model initiative ahead of schedule — JosephJacks_ · 2026-09-19
- Token superposition pretraining idea: prior experiments were 'pretty much a disaster' — giffmana · 2026-09-19
- Talk on scaling fMRI foundation models: CortexMAE and Brainmarks — humanscotti · 2026-09-19
- Score Centering tackles the root cause of RL instability when train and sample policies differ — _AndrewZhao · 2026-09-19
- SELF-INDEX: a framework for retrieval indexes that self-evolve without humans — Sangam Lee · 2026-09-19