AutoDataBench: agents can write quality training tasks but score under 20/100 in 45 minutes

Haotian Luo · hf · 2026-09-30

AutoDataBench evaluates whether agents can synthesize verifiable agentic tasks meeting data-pipeline acceptance criteria (validity, novelty, difficulty, behavioral coverage) — no training run needed. Across three executable-task benchmarks, no agent scores above 20/100 within a 45-minute budget; 4x time helps substantially while per-usable-task cost stays flat. The authors frame autonomous task creation as a key step toward recursive self-improvement. Code and data released.

Original post →

More from Research

Research channel →