Grok 4.5 scores 85.7% on ARC-AGI-1 but just 0.3% on ARC-AGI-3
scaling01 · x · 2026-07-24
Grok 4.5 posts strong ARC-AGI scores, but still collapses on ARC-AGI-3
ARC Prize’s verified results for Grok 4.5 show a sharp split across benchmarks:
- ARC-AGI-1: 85.7% at $0.33/task
- ARC-AGI-2: 52.6% at $0.78/task
- ARC-AGI-3: 0.3% at $6.9K/task
The tester also said raising reasoning effort from medium to high did not improve performance or cost.
More from Models
- Merge launches Fusion, claiming frontier quality at one quarter of the cost — thisguyknowsai · 2026-07-24
- A technical AI crash course covers LLMs, MCP, agents, skills, and RAG — aakashgupta · 2026-07-24
- Gemini 3.5 Flash can build a faithful Minecraft clone — majidmanzarpour · 2026-07-24
- ChatGPT’s public checkout config exposes a new Business ProLite plan — btibor91 · 2026-07-24
- GLM-5.2’s blog hints Z.ai dropped GRPO and went back to PPO — bycloud · 2026-07-24
- Estimating 800B Active Param Model: 10T Total Params if GPT-4 Style — _xjdr · 2026-07-24