ARC-AGI-3, built to resist LLM scaling, reportedly saturated by frontier model Astra

mattturck · x · 2026-09-04

Matt Turck highlights a striking benchmark result: ARC was designed to resist the LLM scaling paradigm. o1 scored just 18% on ARC-AGI in 2024 despite early reasoning ability; the harder ARC-AGI-3 launched in 2026 with frontier AI at only 0.5%; and now Astra has reportedly completely saturated the benchmark using its native harness. If confirmed, it means the benchmark purpose-built to test reasoning and resist scaling has been cracked by the latest frontier model — though the exact harness methodology awaits formal confirmation.

Related event: GPT-6 Astra Reportedly Saturates ARC-AGI-3 with Fewer Steps Than Humans(5 posts)→

Original post →

More from Models

Models channel →