Autonomous Agents Beat Human Baseline in nanoGPT Speedrun Experiment

eliebakouch · x · 2026-08-16

Prime Intellect conducted a massive autonomous agent research experiment using Codex (GPT 5.5) and Claude Code (Opus 4.7) to optimize the nanoGPT training speedrun. After 10k runs and 14k H200 hours, Opus set a new record of 2930 steps, beating the human baseline of 2990. The study reveals that agents excel at hyperparameter sweeps but struggle with novel ideas, and traces the breakdowns in autonomy. All data has been open-sourced.

Related event: Prime Intellect's largest autonomous AI research experiment: 18 models, 153 runs, Fable 5 on top(6 posts)→

Original post →

More from coding & agent

coding & agent channel →