Models rapidly improve at predicting experiment outcomes, may close 90% of gap by 2030

SaxenaNayan · x · 2026-10-07

In a side benchmark, Periodic Labs asked models to predict experiment scores given descriptions and prior run results. Across 1,173 experiments, performance has risen sharply since GPT-4 Turbo, with the best models now closing more than half the gap between a naive baseline and irreducible error — on track to close 90% by October 2030. Notably, no new method or model was invented; the gains came from general capability improvements.

Related event: AI's Ability to Predict Experiments Is Surging, Tests Show(2 posts)→

Original post →

More from Models

Models channel →