GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf

OpenAI's newly released GPT-5.6 Sol demonstrates strong reasoning capabilities across multiple benchmarks, achieving breakthrough results particularly in the ARC-AGI-3 test. It became the first frontier model to successfully crack the ARC-AGI-3 game and topped professional multimodal benchmarks, although it also exposed ongoing shortcomings in handling real-world corporate tasks. Additionally, the model underwent an extended government review period prior to release.

Key Test Data and Details

In the most comprehensive test conducted by the ARC Prize team to date (covering 3 models, 5 inference tiers, 3 benchmark suites, and public/private versions for 90 total slices), GPT-5.6 Sol scored 7.8% on ARC-AGI-3 (some posts note 7.78% or roughly 8%). Compared to the second-place Opus 4.8 at 1.5%, this represents a lead of 500% to 1000%. It also scored 92.5% on ARC-AGI-2, with inference costs an order of magnitude lower than the GPT-5.5 Pro released 3 months ago. GregKamradt noted that the model's performance varies greatly across different tasks, scoring well on levels like ar25 and ft09, but hitting 0% on the hardest levels like g50t, sk48, and su15.

Multimodal Benchmarks and Enterprise Shortcomings

According to @echen, OpenAI referenced the professional multimodal reasoning benchmark GDP.pdf in the GPT-5.6 model card. This dataset contains 100 tasks from real enterprise workflows, written and scored by business experts like doctors and lawyers. GPT-5.6 Sol ranked first with 30.7% accuracy, making it the first model to exceed 30%. However, the model still shows significant flaws when processing mundane documents like insurance claims, incorrectly mapping disaster areas and missing obvious damage while generating authoritative-looking tables. This indicates that AI still needs time to accurately master basic professional tasks.

Release Background and Review

GregKamradt mentioned that GPT-5.6 underwent a US government evaluation period about 4 times longer than usual before release. This extended review, likely required by the government, delayed the launch but resulted in a more stable model checkpoint and more thorough testing analysis.

2026-07-10 ~ 2026-07-11 · 16 related posts

2 near-duplicate retellings: soumitrashukla9 · mhmazur