GPT-6 Astra hits 66% on ARC-AGI-3, near-100% with custom harness at ~$360 per game
AccBalanced · x · 2026-09-05
François Chollet revealed GPT-6 Astra's results on the interactive reasoning benchmark ARC-AGI-3, calling it a step-function change:
- 66% on the standard harness; nearly 100% with a continuous-conversation harness and custom compaction, at roughly $360 per game.
- The continuous-harness version beats the human baseline in action efficiency on almost all levels.
- Inspecting reasoning chains, the team found the model doing efficient on-the-fly symbolic world modeling per game, even inventing its own shorthand DSL—a game-specific algebra.
- ARC's takeaway: capabilities that previously lived in sophisticated harnesses are shifting into the model itself.
A landmark result for interactive reasoning.
Related event: GPT-6 Astra Dominates ARC-AGI-3, Near-Perfect with Custom Harness(4 posts)→
More from Models
- GPT-6 Astra Early Impressions: Reddit Users Call It the Most Capable Model Yet — imadade · 2026-09-05
- GPT-6 Astra's computer use wows users: clicks multiple micro buttons simultaneously — JasonBotterill · 2026-09-05
- OpenAI's early Astra rollout sparks claims it moved to cover up a discovered agent swarm — repligate · 2026-09-05
- New Artificial Analysis Scores Drop, But Qwen 3.8 27B Still Holds Up — RedditUsr2 · 2026-09-05
- 'Massively Disappointed': User Says Astra's Writing Is Only GPT-4 Level, Not AGI — jtteop · 2026-09-05
- Simon Willison benchmarks GPT-6 Astra vs GPT-5.6 on SVG: max effort costs 63 cents per call — gaganghotra_ · 2026-09-05