MazeBench: Gemini 3.8 Flash scores 4%, Fable 5.1 matching its predecessor
patience_cave · x · 2026-09-04
This post is the quoted context of the MazeBench results thread: Gemini 3.8 Flash scored 4%, Muse Spark 1.3 scored 0%, Fable 5.1 is still running, and without code execution models score under 1% in this 3D open-world environment.
Related event: Gemini Flash jumps from 0% to 4% on MazeBench in two months(2 posts)→
More from Models
- New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi — maximelabonne · 2026-09-04
- Sakana AI's Takuya Akiba to unpack Kimi K3's architecture: how a 2.8T-param open model was built — tkasasagi · 2026-09-04
- Small model Luna praised for beating DeepSeek and its uptime for personal agents — bindureddy · 2026-09-04
- GLM-5.3 gets updated chat template: tool-result reordering now exits early — victormustar · 2026-09-04
- Qwopus 3.8 27B Flash fine-tune ships: 12.8% faster decoding, 80.7% MTP acceptance on Qwen3.8-27B — EAccelerate_42 · 2026-09-04
- Gemini 3.8 Flash edges out Astra on DeepSWE: 73.8% vs 73.3% — jon_barron · 2026-09-04