GPT-6 'Astra' burns 60+ hours in 3D spatial reasoning eval, scores just 14%
patience_cave · x · 2026-09-07
Independent eval MazeBench tested "GPT-6 Astra" in a 3D open-world spatial reasoning benchmark. Astra spent over 60 hours in the environment and still finished with only 14% — the best result on the eval so far, but far from solving its hardest challenges.
More from Models
- Google's Astra Agent Allegedly Crushes Existing CAPTCHAs, Sparking Rethink of Bot Checks — eyishazyer · 2026-09-07
- Intelligence price collapsed ~1000x in 18 months as models keep getting smarter — ccerrato147 · 2026-09-07
- Blogger says Google Astra's natural tone broke his last dependency on Claude — StewartalsopIII · 2026-09-07
- VC slams frontier lab subscriptions: usage limits to worsen, top models to shrink — StewartalsopIII · 2026-09-07
- GPT Astra scores just 13% on MazeBench without tools — Wonderful_Buffalo_32 · 2026-09-07
- GPT-6 Astra Generates Stunning Math Animation in One Shot — omarsar0 · 2026-09-07