NanoGPT Speedrun Frontier: 153 Autonomous Runs Across 18 Models Open-Sourced

samsja19 · x · 2026-08-16

PrimeIntellect released full results and traces from the NanoGPT optimizer speedrun benchmark. The evaluation covers 153 autonomous runs across 18 frontier models, including GPT-5.6, Claude, DeepSeek V4, and Kimi K3. The leaderboard tracks best validation loss achieved within 24 hours, percentage of closed test cases, and execution trajectories. All tool calls and generated files are open-sourced for reproducibility.

Related event: Largest Autonomous AI Research Experiment Released(7 posts)→

Original post →

More from coding & agent

coding & agent channel →