alphaXiv's Invent-a-Dataset generates training data from zero seeds, beating frontier APIs on quality and diversity
sarahookr · x · 2026-10-01
Sara Hooker's team at alphaXiv released the technical report for Invent-a-Dataset, a prompt-based system targeting the hardest data regime: building a training set for a capability when you have no seed data at all.
- Input is a natural-language dataset description; output is a large-scale, training-ready post-training dataset
- Benchmarked against five frontier model APIs (Anthropic, Google, OpenAI, DeepSeek, Zhipu) across eight task types and up to 20K samples
- 17% relative quality gains and 19% diversity gains over baselines; the diversity edge widens with scale, from parity at 200 samples to 37% at 20K
- Models fine-tuned on Invent-generated data consistently rank higher than those trained on other generators
A practical recipe for teams facing the cold-start data problem.
More from Research
- Ofir Press: a good benchmark needs scalable data collection — the hardest step yet — OfirPress · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02
- How Matei Zaharia went from Berkeley PhD and Spark to a $190B Databricks — jfiance · 2026-10-02