Grok 4.7 enters AutoResearchExam live leaderboard, ranks No.3 at 30min and No.4 after 24h auto-research
AlexGDimakis · x · 2026-09-24
Grok 4.7 has been added to Bespoke Labs' live AutoResearchExam leaderboard: it jumps to No.3 at 30 minutes of auto-research, drops to No.4 at 4 hours, and lands at No.4 after the full 24-hour run. The author calls its performance/cost tradeoff "pretty good."
AutoResearchExam measures how well agents improve open-ended ML research tasks over 24 hours of auto-research, using hidden-test AUARC (Area Under the AutoResearch Curve) to combine solution quality, generalization, and speed, across 29 tasks. Key design points: one-shot evals miss iterative research dynamics, rankings shift with the evaluation time budget, and the benchmark also compares cost efficiency and how agents allocate research effort.
More from Models
- Qoder offers Qwen3.8-Flash free with no Credits through Sept 30 — Aiden_Tech_Ai · 2026-09-24
- Box says Opus 5.5 cuts token usage 63% and runs 30% faster than Opus 5 — bcherny · 2026-09-24
- Claude Opus 5.5 Max Tops Code Arena WebDev at 1818, Leading GPT-6 Astra by 26 — airesearch12 · 2026-09-24
- 36B Agent Model Runs Fully Local on Snapdragon X2 Elite With Just 32GB RAM — Kyrannio · 2026-09-24
- Opus 5.5 keeps trying to run rm -f; users must repeatedly beg it not to — gandamu_ml · 2026-09-24
- Third-party benchmark: AssemblyAI Universal 3.5 Pro tops 15 STT models at 1.93% WER and 489ms latency — AssemblyAI · 2026-09-24