AutoDataBench: agents can write quality training tasks but score under 20/100 in 45 minutes
Haotian Luo · hf · 2026-09-30
AutoDataBench evaluates whether agents can synthesize verifiable agentic tasks meeting data-pipeline acceptance criteria (validity, novelty, difficulty, behavioral coverage) — no training run needed. Across three executable-task benchmarks, no agent scores above 20/100 within a 45-minute budget; 4x time helps substantially while per-usable-task cost stays flat. The authors frame autonomous task creation as a key step toward recursive self-improvement. Code and data released.
More from Research
- Researcher predicts AI labs will soon pivot from math conjectures to materials and drug discovery — tak3sh8 · 2026-09-30
- Video lecture series by Stephen Wright, Yousef Saad and Peter Bartlett on ML optimization now available — caglar_ee · 2026-09-30
- Arbor: open-source framework for AI agents doing autonomous long-horizon research — burkov · 2026-09-30
- Physics-aware losses keep grain boundaries real when AI generates alloy microstructures — bravo_abad · 2026-09-30
- SOSP26 Paper YoloFS Targets Agent Filesystem Misuse, Built From 290 Real Incident Reports — tianyin_xu · 2026-09-30
- Agents can delete their own logs: Claude Code, Codex, others fail trace integrity, paper finds — maksym_andr · 2026-09-30