MazeBench 3D Spatial Reasoning: Opus 5.5 Hits 6% While GPT-6 Sol and Grok 4.7 Manage Just 1%
patience_cave · x · 2026-09-23
A tough new 3D spatial reasoning benchmark, MazeBench, humbles frontier models: Opus 5.5 leads with a 6% score, inching closer to Astra, while GPT-6 Sol and Grok 4.7 both score only 1%. The results underline how far top models remain from competent 3D spatial reasoning.
More from Models
- NVIDIA releases Nemotron 3 Diarization model handling up to 8 overlapping speakers with 100M params — NVIDIAAI · 2026-09-23
- Dev slams Anthropic's Opus 5.5 safety checks for flagging basic code reviews — evilsocket · 2026-09-23
- Deep conversations with frontier models turn into incomprehensible AI-to-AI jargon, observer warns — erikphoel · 2026-09-23
- Andrew Carr: Opus 5.5 may be the first model that's a bit creative — andrew_n_carr · 2026-09-23
- Anthropic launches Life Sciences Verification Program to gate Opus 5.5 bio access — _sholtodouglas · 2026-09-23
- Rumor: OpenAI's GPT-6 'Astra Minor' is the new Sol, replacing Terra — daniel_mac8 · 2026-09-23