MazeBench 3D Spatial Benchmark: Opus 5.5 Hits 6% While GPT-6 Sol and Grok 4.7 Score Just 1%
daniel_mac8 · x · 2026-09-23
A thread discusses scores on MazeBench, a tough 3D spatial reasoning benchmark: Opus 5.5 scored 6%, inching closer to Astra, while GPT-6 Sol and Grok 4.7 both managed only 1%. The poster also calls a claimed '3x better than Fable 5.1' result interesting. Model and benchmark names are unverified; treat with caution.
More from Models
- NVIDIA releases Nemotron 3 Diarization model handling up to 8 overlapping speakers with 100M params — NVIDIAAI · 2026-09-23
- Dev slams Anthropic's Opus 5.5 safety checks for flagging basic code reviews — evilsocket · 2026-09-23
- Deep conversations with frontier models turn into incomprehensible AI-to-AI jargon, observer warns — erikphoel · 2026-09-23
- Andrew Carr: Opus 5.5 may be the first model that's a bit creative — andrew_n_carr · 2026-09-23
- Anthropic launches Life Sciences Verification Program to gate Opus 5.5 bio access — _sholtodouglas · 2026-09-23
- Rumor: OpenAI's GPT-6 'Astra Minor' is the new Sol, replacing Terra — daniel_mac8 · 2026-09-23