Frontier Model AI Research Experiment: Best Runs Close 82% Gap to Human Records

eliebakouch · x · 2026-08-16

PrimeIntellect ran the largest open experiment evaluating frontier models' ability to conduct AI research. It involved 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. The best runs closed 82% of the gap to a record built by dozens of humans over months.

Related event: Largest Autonomous AI Research Experiment Released(7 posts)→

Original post →

More from coding & agent

coding & agent channel →