Agents now recover 80% of human NanoGPT speedrun gap, up from 40% a year ago

MinqiJiang · x · 2026-08-16

Minqi Jiang notes that on the automated NanoGPT speedrun benchmark, automated agents have closed half of the remaining gap to humans in just over a year. His takeaway: agentic systems will eventually go superhuman at climbing any well-defined metric (e.g. speed, validation loss) on well-defined problems like the nanogpt speedrun — the big open question is whether they can do the same for ill-defined, open-ended problems.

The quoted tweet from benchmark author Bingchen Zhao's team adds context: last year, roughly 6,000 runs with the then-frontier model showed LLMs struggled to make novel algorithmic improvements work, with a best score of 40% of the gap to the human record recovered; this year that number has been pushed to 80% by Fable 5, and agents may genuinely surpass humans next year. The underlying paper, "The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements" (arXiv:2506.22419), builds 19 tasks from the community NanoGPT speedrun, giving the agent the previous record's training script plus optional hints in three formats ranging from pseudocode to paper-like descriptions, to test whether agents can reproduce improvements in an active research area.

Original post →

More from coding & agent

coding & agent channel →