OpenAI Responds to Benchmark Concerns: GDPval Nearing Saturation
emollick · x · 2026-07-31
OpenAI officially responded to recent benchmark dynamics. For ARC-AGI-3, human testers scored around 48%.
Regarding GDPval, OpenAI stated that the benchmark is now close to saturated, prompting a shift in focus to other evaluation methods. While GDPval was a great eval, its tasks were much more heavily specified than real-world usage. Currently unsaturated benchmarks include ARC-AGI, the original GDPval, METR long horizons, and ASI cyber tasks.
More from Models
- Grok 2 Voice Mode Delayed? X Platform Silently Changes Release Window to 'Soon' — teortaxesTex · 2026-07-31
- Anthropic Launches Claude Opus 5: State-of-the-Art Performance at Half the Cost — dl_weekly · 2026-07-31
- Osmosis AI Bets on RL and YC GPU Cluster to Rival Foundation Models — togethercompute · 2026-07-31
- ChatGPT Rate Limit Quirk: No Self-Check and Independent of Codex — Sauers_ · 2026-07-31
- Annotator Ghost in the Weights: Opus 5 Channels Trainers in Base Mode — aiamblichus · 2026-07-31
- Google Rollsouts Gemini Robotics ER 2 Model via API — ericjang11 · 2026-07-31