GPT-5.6 Sol's ARC-AGI-3 Difficulty Perception

GregKamradt · x · 2026-07-10

The post suggests sorting by GPT-5.6 Sol's scores to grasp the difficulty of the public ARC-AGI-3 tasks. By sharing the scores of the easiest and hardest games, the author illustrates a massive variance between problems, which indirectly reflects the model's performance distribution on this benchmark.

Related event: GPT-5.6 Sets New SOTA on ARC-AGI-3 and Exceeds 30% on GDP.pdf(16 posts)→

Original post →

More from Models

Models channel →