AutoResearchExam: 29-task benchmark ranks agents on 24-hour open-ended ML research
AlexGDimakis · x · 2026-09-25
Bespoke Labs (with Dimitris Dimakis) released AutoResearchExam, a benchmark measuring how well agents sustain progress on open-ended ML research tasks.
Key points:
- 29 tasks, 24-hour auto-research windows: the metric, hidden-test AUARC (Area Under the AutoResearch Curve), combines solution quality, generalization to a hidden test set, and speed.
- Separates visible progress from overfitting: agents optimize validation scores, but rankings are set by a hidden test set, avoiding mistaking validation overfitting for real progress.
- Rankings shift with time budget: the analysis shows leaderboards change depending on the evaluation window, so a single budget can hide model differences; it also compares cost efficiency and how agents allocate research effort.
- Opus 5.5 is the new leaderboard leader and resets the performance-vs-cost Pareto frontier, pushing Fable 5.1 and Astra 6 off the frontier.
Related event: Bespoke Labs Launches AutoResearchExam Benchmark for Research Agents(2 posts)→
More from Models
- AIs hit the maximum 151 IQ on Mensa Norway test as Musk flags the trend — elonmusk · 2026-09-25
- GPT-Sol's Chinese output bleed isn't heavy quantization—it's language confusion, researcher says — AryHHAry · 2026-09-25
- TypeSafe's Jev model explained: calibrated confidence, constrained outputs, 'zero hallucinations' — tech_technical · 2026-09-25
- OpenAI to preview GPT-6 Cyber model and first-of-its-kind security product, per Fortune — jeremyakahn · 2026-09-25
- Vision model tier list updated with Opus 5.5, GPT-6 Sol/Luna, and Grok 4.7 — ducha_aiki · 2026-09-25
- Anthropic Accused of Quietly Nerfing Models Weeks After Launch, Opus 5.5 Expected to Follow — iannuttall · 2026-09-25