GPT-5.6 Leads Professional Multimodal Benchmark GDP.pdf
echen · x · 2026-07-11
OpenAI and Anthropic have recently cited a professional multimodal reasoning benchmark known as GDP.pdf. This benchmark focuses on mundane professional documents like contracts and clinical notes.
In the latest tests, GPT-5.6 Sol took first place with an accuracy of 30.7%, making it the first model to break the 30% barrier.
Related event: GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf(16 posts)→
More from Models
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22