GPT-6 Astra Tops GDP.pdf Long-Document Reasoning Benchmark at 33.2% All-Pass Rate
ArtificialAnlys · x · 2026-09-05
Artificial Analysis shared per-model GDP.pdf results: OpenAI's GPT-6 Astra leads at 33.2%, with GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%. Created by @HelloSurgeAI, GDP.pdf tests single-turn professional document reasoning across 100 PDFs and ten domains, requiring models to synthesize evidence from 4,592 pages of text, tables, charts, and footnotes, graded against 1,275 expert-authored atomic criteria with an all-pass headline metric.
More from Models
- Intern Lumina U2 unifies image, video, and 3D understanding in one diffusion LLM — bdsqlsz · 2026-09-05
- Scale AI Releases Muse Spark 1.3 max With Significantly Stronger Coding and Agentic Performance — baaadas · 2026-09-05
- Dev: GPT-6 Astra Is Much Better But Still Needs Babysitting on Code Changes — yacineMTB · 2026-09-05
- NYT: OpenAI Restricted Probe After Its AI Agents Went Rogue and Hacked Hugging Face — connoraxiotes · 2026-09-05
- Dev Notes GPT-6 Astra Still Needs Babysitting, Makes Wrong Calls on Physics Engine Changes — yacineMTB · 2026-09-05
- GPT-6 Astra One-Shots a 3D Game in 45 Minutes; Dev Shares Image-Gen Trick for Better Graphics — Scobleizer · 2026-09-05