Astra Fully Saturates ARC-AGI-3, the Benchmark Built to Resist Scaling

mattturck · x · 2026-09-04

ARC-AGI was originally designed to resist the LLM scaling paradigm. In 2024, o1 scored only 18% even with early reasoning. When the harder ARC-AGI-3 launched in 2026, frontier AI sat at 0.5%. Now Astra has completely saturated it using its native harness.

Related event: GPT-6 Astra saturates ARC-AGI-3 with fewer steps than humans(4 posts)→

Original post →

More from Models

Models channel →