Bespoke Labs Launches AutoResearchExam: Benchmarking 24-Hour Auto-Research Agents
AlexGDimakis · x · 2026-09-24
Bespoke Labs released AutoResearchExam, a benchmark of 29 open-ended ML research tasks measuring how fast agents improve a private test score over 24-hour auto-research. Its hidden-test AUARC metric separates validation progress from true generalization; rankings shift with the time budget, and the analysis covers cost efficiency and agent research behavior. Endorsed by Emad Mostaque.
More from coding & agent
- The minimal agent stack: Builder, Reviewer and Verifier in an escalation loop — BLUECOW009 · 2026-09-24
- Claude Opus 5.5 tops the Coding Agent Index at 66, but cost per task jumps 21% — ArtificialAnlys · 2026-09-24
- +$2M ARR after the tweet: customers build their own Claude workflows instead of buying tools — garrytan · 2026-09-24
- Qwen Image 2.1 masked inpainting: working crop-and-stitch graph, wiring and prompting differences from Flux — Reasonable_Arm7239 · 2026-09-24
- Jev Engineering: use a tiny model as your agent's brain to slash token bills — blaizedsouza · 2026-09-24
- PatronusAI evals panel covers SpeedrunBench, world models, and spreadsheet agent evals — DynamicWebPaige · 2026-09-24