MazeBench results: Gemini 3.8 Flash scores 4%, most models under 1% in 3D open world

patience_cave · x · 2026-09-04

patiencecave published new MazeBench results, a benchmark testing visual-spatial reasoning in a 3D open-world maze: Gemini 3.8 Flash scored 4%, Muse Spark 1.3 scored 0%, and Fable 5.1 is still running, currently matching Fable 5.

Key takeaway: without code execution access, models score under 1% in this environment. Follow-ups: GPT-5.6 Sol holds the top score thanks to a longer, cheaper run; Fable 5.1 has reached 10% and could take the lead; Google's Flash line went 0% → 1% → 4% in a couple of months, with heat maps showing each new agent models the world more accurately.

Original post →

More from Models

Models channel →