18 models compete in autonomous coding: top closes 82% human gap
AI寒武纪 · wechat · 2026-08-16
PrimeIntellect conducted the largest-scale autonomous AI research experiment to date, placing 18 top models in an isolated sandbox to optimize nanoGPT.
Key Findings:
- Clear Stratification: Fable5 achieved 2,726 steps, closing 81.7% of the gap to human records. Opus5 and KimiK3 followed closely, while others like GPT-5.5 performed near-randomly.
- Gap Stems from Research Taste: Stronger models (e.g., Opus5) effectively model noise and re-test failed hypotheses with ablation, whereas weaker models abandon promising directions based on single failures.
- Tool-Building Capabilities: Some models built internal numerical labs or utilized multi-agent frameworks (theorist, auditor) to validate hypotheses in low-cost synthetic environments before committing to expensive training runs.
- Lack of Novelty: Despite strong execution and understanding, no model produced fundamentally novel methods, relying instead on combinations and tuning of known techniques.
Related event: Prime Intellect Releases Largest Autonomous AI Research Experiment(9 posts)→
More from coding & agent
- Dev showcases running 48 AI agents in parallel — TejasKumar_ · 2026-08-16
- How to build shared context between WhatsApp and AI voice agents — Madhav_Agarwal_ · 2026-08-16
- Sunil Pai's experiment: LLM-enriched voice transcription that links issues and wiki live — threepointone · 2026-08-16
- Do we need gateways now that MCPs have gone stateless? — wallphaser231 · 2026-08-16
- Codex intelligently routes tasks to cheaper Luna model to save limits — eyishazyer · 2026-08-16
- SelfMem Paper: AI Agents That Manage Their Own Memory Outperform Fixed Systems — alex_verem · 2026-08-16