ARC-AGI-3, built to resist LLM scaling, reportedly saturated by frontier model Astra
mattturck · x · 2026-09-04
Matt Turck highlights a striking benchmark result: ARC was designed to resist the LLM scaling paradigm. o1 scored just 18% on ARC-AGI in 2024 despite early reasoning ability; the harder ARC-AGI-3 launched in 2026 with frontier AI at only 0.5%; and now Astra has reportedly completely saturated the benchmark using its native harness. If confirmed, it means the benchmark purpose-built to test reasoning and resist scaling has been cracked by the latest frontier model — though the exact harness methodology awaits formal confirmation.
Related event: GPT-6 Astra Reportedly Saturates ARC-AGI-3 with Fewer Steps Than Humans(5 posts)→
More from Models
- GPT-6 Astra hits 66% on ARC-AGI-3, up from 8%, as ARC Prize plans AGI-4 around open-ended invention — GaryMarcus · 2026-09-04
- Scaling Law for Looped Transformers: Looping Boosts Reasoning, Not Knowledge — bookwormengr · 2026-09-04
- AI experts once pegged AGI at 2075-2100 — OpenAI's Astra shows how wrong they were — dee_hw · 2026-09-04
- Models still can't see determinants, but no-thinking time-horizons double every ~9 months — scaling01 · 2026-09-04
- GPT-6 Astra Ultra builds a Minecraft wooden house zero-shot in ~9 minutes, stairs perfect — Angaisb_ · 2026-09-04
- Benchmark saturation accelerates: ARC-AGI-3 falls in 5 months after ARC-AGI-1's 6 years — mmmbchang · 2026-09-04