OpenAI Responds to Benchmark Concerns: GDPval Nearing Saturation

emollick · x · 2026-07-31

OpenAI officially responded to recent benchmark dynamics. For ARC-AGI-3, human testers scored around 48%.

Regarding GDPval, OpenAI stated that the benchmark is now close to saturated, prompting a shift in focus to other evaluation methods. While GDPval was a great eval, its tasks were much more heavily specified than real-world usage. Currently unsaturated benchmarks include ARC-AGI, the original GDPval, METR long horizons, and ASI cyber tasks.

Original post →

More from Models

Models channel →