Bespoke Labs launches AutoResearchExam: 29 tasks to benchmark 24-hour autonomous ML research agents
AashaySachdeva · x · 2026-09-10
- Bespoke Labs releases AutoResearchExam, a benchmark of 29 open-ended ML research tasks measuring how fast agents improve a private test score over a 24-hour sustained research window.
- Motivation: real research relies on iteration and exploration; one-shot and standard long-horizon evals miss this and can conflate validation-score gains with true generalization.
- Design: agents optimize a visible validation score while a hidden-test AUARC measures how quickly they find generalizing solutions.
- Findings: rankings shift with the evaluation window, showing a single time budget misses model differences; the report also analyzes cost efficiency and how agents allocate research effort.
- Runs on the standard Terminus 2 harness; tasks and leaderboard are public. Framed as an ingredient toward recursive self-improvement.
Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→
More from coding & agent
- Letting GPT-6 Astra drive software UIs directly beats built-in agents, dev reports — brandon_galang · 2026-09-10
- Cost per agent-written PR drops to $30, team targets $10 — vikvang1 · 2026-09-10
- Developer shows off Codex Remote Duo setup, drawing attention — Dimillian · 2026-09-10
- DeepLearning.AI maps the core skills for steering coding agents in its AI engineering framework — DeepLearningAI · 2026-09-10
- The Dockerfile agent also detects required secrets—and the UX is surprisingly hard — lucasmeijer · 2026-09-10
- Agent auto-builds a custom Dockerfile to optimize new workspace start times — lucasmeijer · 2026-09-10