Models rapidly improve at predicting experiment outcomes, may close 90% of gap by 2030
SaxenaNayan · x · 2026-10-07
In a side benchmark, Periodic Labs asked models to predict experiment scores given descriptions and prior run results. Across 1,173 experiments, performance has risen sharply since GPT-4 Turbo, with the best models now closing more than half the gap between a naive baseline and irreducible error — on track to close 90% by October 2030. Notably, no new method or model was invented; the gains came from general capability improvements.
Related event: AI's Ability to Predict Experiments Is Surging, Tests Show(2 posts)→
More from Models
- vLLM adds day-0 support for Google's multimodal EmbeddingGemma 2 — AccBalanced · 2026-10-07
- RedNote Has a Serious, Low-Key AI Lab, Argues Analyst Weighing Dots3 Results — teortaxesTex · 2026-10-07
- OpenAI Cuts API Rate Limit Tiers From Five to Three, Grow Tier Now at $500 — OpenAIDevs · 2026-10-07
- Users report GPT-6.1 Sol is slower and worse in practice, suspect sandbagging — rickasaurus · 2026-10-07
- 61–90% of AI-synthesized medical conclusions contain factual errors, SciConBench finds — manoelribeiro · 2026-10-07
- Cohere Labs Debuts Tiny Aya, a Small-Model Family Covering 70+ Languages — Cohere_Labs · 2026-10-07