MazeBench results: Gemini 3.8 Flash scores 4%, most models under 1% in 3D open world
patience_cave · x · 2026-09-04
patiencecave published new MazeBench results, a benchmark testing visual-spatial reasoning in a 3D open-world maze: Gemini 3.8 Flash scored 4%, Muse Spark 1.3 scored 0%, and Fable 5.1 is still running, currently matching Fable 5.
Key takeaway: without code execution access, models score under 1% in this environment. Follow-ups: GPT-5.6 Sol holds the top score thanks to a longer, cheaper run; Fable 5.1 has reached 10% and could take the lead; Google's Flash line went 0% → 1% → 4% in a couple of months, with heat maps showing each new agent models the world more accurately.
More from Models
- New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi — maximelabonne · 2026-09-04
- Sakana AI's Takuya Akiba to unpack Kimi K3's architecture: how a 2.8T-param open model was built — tkasasagi · 2026-09-04
- Small model Luna praised for beating DeepSeek and its uptime for personal agents — bindureddy · 2026-09-04
- GLM-5.3 gets updated chat template: tool-result reordering now exits early — victormustar · 2026-09-04
- Qwopus 3.8 27B Flash fine-tune ships: 12.8% faster decoding, 80.7% MTP acceptance on Qwen3.8-27B — EAccelerate_42 · 2026-09-04
- Gemini 3.8 Flash edges out Astra on DeepSWE: 73.8% vs 73.3% — jon_barron · 2026-09-04