MazeBench 3D Spatial Reasoning: Opus 5.5 Hits 6% While GPT-6 Sol and Grok 4.7 Manage Just 1%

patience_cave · x · 2026-09-23

A tough new 3D spatial reasoning benchmark, MazeBench, humbles frontier models: Opus 5.5 leads with a 6% score, inching closer to Astra, while GPT-6 Sol and Grok 4.7 both score only 1%. The results underline how far top models remain from competent 3D spatial reasoning.

Original post →

More from Models

Models channel →