Autonomous Agents Beat Human Baseline in nanoGPT Speedrun Experiment
eliebakouch · x · 2026-08-16
Prime Intellect conducted a massive autonomous agent research experiment using Codex (GPT 5.5) and Claude Code (Opus 4.7) to optimize the nanoGPT training speedrun. After 10k runs and 14k H200 hours, Opus set a new record of 2930 steps, beating the human baseline of 2990. The study reveals that agents excel at hyperparameter sweeps but struggle with novel ideas, and traces the breakdowns in autonomy. All data has been open-sourced.
More from coding & agent
- VT Code 0.146.0 updates: adds Gemini 3.7 Flash and Qwen3.8 27B support — shensi · 2026-08-16
- How tech leaders actually use AI agents in their day to day workflow — shensi · 2026-08-16
- Anthropic's new Computer Use tool praised for speed and usability — Daniel_Farinax · 2026-08-16
- llama.cpp Windows Manager: Open-source visual tool for managing multiple models — wgaca2 · 2026-08-16
- Use Fast Models for Interaction, Slow for Background: Why Grok 4.6 Fits — vikvang1 · 2026-08-16
- AFK Pilot Relay Open Sourced: Secure Message Routing for Coding Agents — PawelHuryn · 2026-08-16