GPT-5.6 Cites Multimodal Benchmark and Breaks 30%

echen · x · 2026-07-11

OpenAI referenced GDP.pdf in the GPT-5.6 model card released yesterday, which is a professional multimodal reasoning benchmark from Surge.

The poster noted that OpenAI isn't the first frontier lab to cite it this year; Anthropic's Fable 5 previously referenced this benchmark. Meanwhile, GPT-5.6 Sol currently leads with 30.7%, becoming the first model to break the 30% mark.

Related event: GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf(16 posts)→

Original post →

More from Models

Models channel →