Astra Fully Saturates ARC-AGI-3, the Benchmark Built to Resist Scaling
mattturck · x · 2026-09-04
ARC-AGI was originally designed to resist the LLM scaling paradigm. In 2024, o1 scored only 18% even with early reasoning. When the harder ARC-AGI-3 launched in 2026, frontier AI sat at 0.5%. Now Astra has completely saturated it using its native harness.
Related event: GPT-6 Astra saturates ARC-AGI-3 with fewer steps than humans(4 posts)→
More from Models
- Sean Taylor claims fast progress eradicating hallucinations; Andrew Ng: capabilities and safety can align — irinarish · 2026-09-04
- ValsAI says OpenAI's GPT 6 Astra has effectively saturated SRE-Bench reverse-engineering benchmark — sandersted · 2026-09-04
- OpenAI launches GPT-6 Astra, its first cyber-critical model under Preparedness Framework — alishbaimran_ · 2026-09-04
- Alignment researcher: Astra's near-zero misalignment looks like whack-a-mole, not real fix — sjgadler · 2026-09-04
- Google confirms Gemini 3.8 Flash in AI Mode drops citations and links, fix on the way — gaganghotra_ · 2026-09-04
- Quick Question: Does GPT-6 Include HuggingFace Access? — gordic_aleksa · 2026-09-04