GPT-6 Astra Clears All 25 ARC-AGI-3 Games, Hitting ~80% of Optimal Play
i_dg23 · x · 2026-09-13
- Mike Knoop discusses the limits of RSI: intelligence isn't an unbounded scalar; it can be measured as decision quality vs. optimality, capped at 100%. Astra is already 80% optimal on ARC v3 speedruns.
- Context: on the 25 public ARC-AGI-3 games, GPT-6 Astra scored 100% using a new provider adapter harness (preserving opaque reasoning across requests, enabling auto-compaction), compared against human baselines and best-known fewest-action runs.
- Knoop sees two near-term areas where RSI matters: horizontal data acquisition (automating data generation to fill knowledge gaps in weights) and efficiency/cost (far from optimality, ideal for autoresearch — cheaper models enable even more autoresearch).
More from Models
- Opus beats Astra on one-shot 3D quality; Astra is 4x faster and 10x more token-efficient — chaseleantj · 2026-09-13
- Early Muse Spark 1.3 hands-on shows flawed image generation, seems not benchmaxxed — teortaxesTex · 2026-09-13
- User's Hermes setup running DeepSeek launches eerie program and speaks in three voices unattended — Teknium · 2026-09-13
- OpenAI hits automated research intern goal, agents solve Navier-Stokes in 88 hours — btibor91 · 2026-09-13
- GPT-6 Astra beats human drone baseline and earns 3x Claude on Vending-Bench — The Decoder · 2026-09-13
- Local LLM model picker: how to choose between Llama, Mistral, Qwen and DeepSeek — anant94 · 2026-09-13