Bespoke Labs Launches AutoResearchExam: Benchmarking 24-Hour Auto-Research Agents

AlexGDimakis · x · 2026-09-24

Bespoke Labs released AutoResearchExam, a benchmark of 29 open-ended ML research tasks measuring how fast agents improve a private test score over 24-hour auto-research. Its hidden-test AUARC metric separates validation progress from true generalization; rankings shift with the time budget, and the analysis covers cost efficiency and agent research behavior. Endorsed by Emad Mostaque.

Original post →

More from coding & agent

coding & agent channel →