Grok 4.5 scores 85.7% on ARC-AGI-1 but just 0.3% on ARC-AGI-3

scaling01 · x · 2026-07-24

Grok 4.5 posts strong ARC-AGI scores, but still collapses on ARC-AGI-3

ARC Prize’s verified results for Grok 4.5 show a sharp split across benchmarks:

The tester also said raising reasoning effort from medium to high did not improve performance or cost.

Original post →

More from Models

Models channel →