Grok 4.7 enters AutoResearchExam live leaderboard, ranks No.3 at 30min and No.4 after 24h auto-research

AlexGDimakis · x · 2026-09-24

Grok 4.7 has been added to Bespoke Labs' live AutoResearchExam leaderboard: it jumps to No.3 at 30 minutes of auto-research, drops to No.4 at 4 hours, and lands at No.4 after the full 24-hour run. The author calls its performance/cost tradeoff "pretty good."

AutoResearchExam measures how well agents improve open-ended ML research tasks over 24 hours of auto-research, using hidden-test AUARC (Area Under the AutoResearch Curve) to combine solution quality, generalization, and speed, across 29 tasks. Key design points: one-shot evals miss iterative research dynamics, rankings shift with the evaluation time budget, and the benchmark also compares cost efficiency and how agents allocate research effort.

Original post →

More from Models

Models channel →