Astra Fully Saturates the ARC-AGI-3 Benchmark Using Fewer Moves Than the Average Human

ObiWanCanownme · reddit · 2026-09-04

According to ARC Prize's official blog, the Astra agent has saturated the new interactive reasoning benchmark ARC-AGI-3 — hitting the benchmark's ceiling score — and did so using fewer moves on average than human testers, marking one of the strongest results yet recorded on it.

Related event: GPT-6 Astra saturates ARC-AGI-3 with near-zero reasoning tokens(6 posts)→

Original post →

More from Models

Models channel →