MazeBench 3D Spatial Benchmark: Opus 5.5 Hits 6% While GPT-6 Sol and Grok 4.7 Score Just 1%

daniel_mac8 · x · 2026-09-23

A thread discusses scores on MazeBench, a tough 3D spatial reasoning benchmark: Opus 5.5 scored 6%, inching closer to Astra, while GPT-6 Sol and Grok 4.7 both managed only 1%. The poster also calls a claimed '3x better than Fable 5.1' result interesting. Model and benchmark names are unverified; treat with caution.

Original post →

More from Models

Models channel →