GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf
OpenAI's newly released GPT-5.6 Sol demonstrates strong reasoning capabilities across multiple benchmarks, achieving breakthrough results particularly in the ARC-AGI-3 test. It became the first frontier model to successfully crack the ARC-AGI-3 game and topped professional multimodal benchmarks, although it also exposed ongoing shortcomings in handling real-world corporate tasks. Additionally, the model underwent an extended government review period prior to release.
Key Test Data and Details
In the most comprehensive test conducted by the ARC Prize team to date (covering 3 models, 5 inference tiers, 3 benchmark suites, and public/private versions for 90 total slices), GPT-5.6 Sol scored 7.8% on ARC-AGI-3 (some posts note 7.78% or roughly 8%). Compared to the second-place Opus 4.8 at 1.5%, this represents a lead of 500% to 1000%. It also scored 92.5% on ARC-AGI-2, with inference costs an order of magnitude lower than the GPT-5.5 Pro released 3 months ago. GregKamradt noted that the model's performance varies greatly across different tasks, scoring well on levels like ar25 and ft09, but hitting 0% on the hardest levels like g50t, sk48, and su15.
Multimodal Benchmarks and Enterprise Shortcomings
According to @echen, OpenAI referenced the professional multimodal reasoning benchmark GDP.pdf in the GPT-5.6 model card. This dataset contains 100 tasks from real enterprise workflows, written and scored by business experts like doctors and lawyers. GPT-5.6 Sol ranked first with 30.7% accuracy, making it the first model to exceed 30%. However, the model still shows significant flaws when processing mundane documents like insurance claims, incorrectly mapping disaster areas and missing obvious damage while generating authoritative-looking tables. This indicates that AI still needs time to accurately master basic professional tasks.
Release Background and Review
GregKamradt mentioned that GPT-5.6 underwent a US government evaluation period about 4 times longer than usual before release. This extended review, likely required by the government, delayed the launch but resulted in a more stable model checkpoint and more thorough testing analysis.
2026-07-10 ~ 2026-07-11 · 16 related posts
- GPT-5.6-Sol Scores Higher on ARC-AGI-3 — Scobleizer · 2026-07-10
- GPT-5.6 Sol Sets New Record on ARC-AGI-3 — scaling01 · 2026-07-10
- [source] OpenAI Scores High on ARC-AGI 3 — ChrissGPT · 2026-07-10
- ChatGPT 5.6 Scores on ARC-AGI 3 — Bizzyguy · 2026-07-10
- GPT-5.6 Sol Dominates ARC-AGI-3 — haider1 · 2026-07-10
- GPT-5.6 Sol Sets New ARC Record — scaling01 · 2026-07-10
- [source] GPT-5.6 Undergoes Most Comprehensive ARC Evaluation — debashis_dutta · 2026-07-10
- GPT-5.6 Sol's ARC-AGI-3 Difficulty Perception — GregKamradt · 2026-07-10
- ARC-AGI-3 Difficulty Distribution and Model Scores — GregKamradt · 2026-07-10
- GPT-5.6 Sets New Record on ARC-AGI-3, Delayed by Review — soumitrashukla9 · 2026-07-10
- GPT-5.6 Cites Multimodal Benchmark and Breaks 30% — echen · 2026-07-11
- [source] GPT-5.6 Leads Professional Multimodal Benchmark GDP.pdf — echen · 2026-07-11
- SOTA Models Fail at Processing Insurance Claims — echen · 2026-07-11
- GDP Benchmark: Evaluating Real Enterprise Tasks — echen · 2026-07-11
2 near-duplicate retellings: soumitrashukla9 · mhmazur