Frontier Models Close 82% of Human Research Gap in Auto-Research Experiment
eliebakouch · x · 2026-08-16
PrimeIntellect released the largest open experiment on auto-research capabilities of frontier models. With 100+ autonomous runs across 10+ models sandboxed on 8xH200s for up to 8 days, the models iterated on the nanoGPT optimizer track. The best runs closed 82% of the performance gap to a record built by human experts over months.
Related event: Prime Intellect Releases Largest Autonomous AI Research Experiment Results(8 posts)→
More from Research
- CUHK's VideoCoCo: executable code as CoT lifts VBench-2.0 average by 25.7 points — 机器之心 · 2026-08-16
- Open-source proxy NullOrigin strips KGW watermarks from LLM outputs in real-time — theawkwardbong · 2026-08-16
- How AI text watermarking works and how to evade it, as Anthropic adopts it — SpiritRealistic8174 · 2026-08-16
- Viral AI sparks biosecurity debate: sparking a pandemic is easier than defending one — anshulkundaje · 2026-08-16
- ORBIT Training Paradigm Boosts Zero-Shot Forecasting for Time Series Foundation Models — chaumian · 2026-08-16
- Hamilton-Zero: A Neural Foundation Model for Solving Quantum Ground States — burny_tech · 2026-08-16