Astra tested on ~30 obscure puzzle games: ARC-AGI-3's ~99% may undersell it
burny_tech · x · 2026-09-09
Blogger FakePsyho tested Astra's reasoning in a sandboxed setup—no internet, no coding tools—having Astra+codex play around 30 obscure Puzzlescript puzzle games featuring rule discovery and brutally hard levels.
Key observations:
- Astra can do pathfinding and planning without code, often batching 20-30 actions in a row, on par with expert puzzle players.
- Its mistakes are very human-like (e.g., picking longer paths that require less thinking), leading the author to believe there's no code execution under the hood.
- The author argues the 99% score on ARC-AGI-3 might actually undersell it—the agent is already above an average human player.
More from Models
- GPT 6 Astra reportedly one-shots a Re-Volt clone in 20 minutes, then plays it itself — nickbaumann_ · 2026-09-09
- Zuckerberg: Meta already training post-Watermelon models on its 1GW Prometheus cluster — rohanpaul_ai · 2026-09-09
- Dev Pushes Back on Astra Hype: Being Good at Blender Isn't an AGI Benchmark — carsonfarmer · 2026-09-09
- Meta's Muse Spark 1.3 Max lands #8 on Code Arena: WebDev, reshaping the price-performance frontier — arena · 2026-09-09
- Aidan Gomez: labs train on rewritten user data even under ZDR promises — josh_wills · 2026-09-09
- Open problems turned into RL environments: benchmarks and RL envs are two sides of the same coin — burny_tech · 2026-09-09