GPT-5.6 Leads Professional Multimodal Benchmark GDP.pdf

echen · x · 2026-07-11

OpenAI and Anthropic have recently cited a professional multimodal reasoning benchmark known as GDP.pdf. This benchmark focuses on mundane professional documents like contracts and clinical notes.

In the latest tests, GPT-5.6 Sol took first place with an accuracy of 30.7%, making it the first model to break the 30% barrier.

Related event: GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf(16 posts)→

Original post →

More from Models

Models channel →