GPT-6 Sol scores 89.6% on ARC-AGI-2 but only 23% on ARC-AGI-3, ARC Prize reports
fchollet · x · 2026-09-29
ARC Prize published verified ARC-AGI results for OpenAI's GPT-6 Sol:
- ARC-AGI-3: 4.6% ($5.6K) on the standard harness, 23.0% ($8.7K) with the provider adapter harness — far behind Astra's 99.9%, though well above Luna's 0.59%
- ARC-AGI-2: 89.6% at $0.44/task
- ARC-AGI-1: 95.5% at $0.14/task
Sol effectively solves ARC-AGI-2, but the huge gap on ARC-AGI-3 — which tests open-world exploration and reasoning — shows frontier models still struggle to generalize in genuinely novel environments.
More from Models
- ElevenLabs v4 Voice Model Now Available on Runway Platform — runwayml · 2026-09-29
- Anthropic ships Sonnet 5.5 at half the price, nearly matching Opus 5.5; Haiku 5.5 on the way — oran_ge · 2026-09-29
- Xiaomi MiMo-V2.6-Distill-Qwen-9B GGUF quantization trends on Hugging Face — bartowski · 2026-09-29
- OpenAI reportedly cancels October release of GPT-6.1 Astra over safety concerns — Polymarket · 2026-09-29
- Testing 5 Qwen3.6-35B-A3B finetunes: the base model beats almost all of them — returnity · 2026-09-29
- Replay agents hit SOTA on CUA benchmarks: NeurIPS oral paper exposes eval flaws — proceduralia · 2026-09-29