Astra Fully Saturates the ARC-AGI-3 Benchmark Using Fewer Moves Than the Average Human
ObiWanCanownme · reddit · 2026-09-04
According to ARC Prize's official blog, the Astra agent has saturated the new interactive reasoning benchmark ARC-AGI-3 — hitting the benchmark's ceiling score — and did so using fewer moves on average than human testers, marking one of the strongest results yet recorded on it.
Related event: GPT-6 Astra saturates ARC-AGI-3 with near-zero reasoning tokens(6 posts)→
More from Models
- ChatGPT's image generator is much improved via Astra — joshua_saxe · 2026-09-05
- Fable 5.1 vs GPT 6 Astra on 3D Blender asset generation shows a stark gap — curious_capsuleer · 2026-09-05
- GPT 6 'Astra' reportedly recreates Pokémon from a single prompt — IanArawjo · 2026-09-05
- User switches back to GPT from Gemini after 9/10 PDF failures and heavy usage burn — DrlNoV · 2026-09-05
- GPT-6 Astra's experimental compaction in Codex saves notes across context windows, off by default — TheMoonMidas · 2026-09-05
- Same prompt run on Opus 4.8, Fable 5.1, and GPT 6 Astra for comparison — LightningMcLovin · 2026-09-05