Frontier Models Close 82% of Human Research Gap in Auto-Research Experiment

eliebakouch · x · 2026-08-16

PrimeIntellect released the largest open experiment on auto-research capabilities of frontier models. With 100+ autonomous runs across 10+ models sandboxed on 8xH200s for up to 8 days, the models iterated on the nanoGPT optimizer track. The best runs closed 82% of the performance gap to a record built by human experts over months.

Related event: Prime Intellect Releases Largest Autonomous AI Research Experiment Results(8 posts)→

Original post →

More from Research

Research channel →