Adaption's Invent a Dataset generates training-ready datasets with zero seed data, beating GPT and Claude on diversity
sarahookr · x · 2026-10-02
Adaption, with collaborators including Sara Hooker, introduced Invent a Dataset, research into autonomously building training datasets with zero seed data — from a dataset description and requested volume straight to a training-ready dataset.
Key points:
- Dataset construction remains one of the most manual and brittle parts of AI development; the core challenge is diversity collapse as requested volume scales up, plus maintaining generation quality.
- The team benchmarked Invent API against Anthropic, Google, OpenAI, DeepSeek, and Zai APIs plus open-weight models, across tasks, domains, and languages, at 200 to 20,000 samples.
- Invent's datasets are 19%-55% more diverse than every model tested and about 17% higher quality than the strongest baseline (Claude Opus 5), improving on both axes simultaneously (Pareto frontier).
- Unlike some providers that bar training on synthetic outputs, Invent explicitly permits commercial downstream training use.
More from Models
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Google's Argon battle-tested by 200k+ Googlers daily, not benchmaxxed — Zergylord · 2026-10-02
- JEV claims to be the first System One model hosted in the EU — juanviera23 · 2026-10-02
- Gemini knew a user's mom's name unprompted, then gave three conflicting explanations — Exact_Firefighter864 · 2026-10-02
- GPT-6.1 Sol nearly matches Astra at a quarter of the price on RareBench — danielmckinn0n · 2026-10-02
- New results from PostTrainBench v1.2 are in — mariofilhoml · 2026-10-02