MazeBench results: flagship models score 0% on maze tasks
patience_cave · x · 2026-08-24
New MazeBench results are in: ox-alpha, Grok 4.6, GLM-5.3, and Qwen 3.8-max all score 0% on maze-solving tasks; Gemini 3.7 Flash manages 1% and beats Kimi K3. The takeaway: frontier models remain collectively terrible at structured spatial reasoning.
Related event: MazeBench: Most Flagship Models Score Zero on Maze Tasks(3 posts)→
More from Models
- Qwen 3.8 27B ranks 9th on Code Arena — tarruda · 2026-08-25
- Kimi K3 surpasses Sol 5.6 and Opus 4.8 on Code Arena — bgmshana · 2026-08-25
- GLiNER2.5 switches to boundary prediction, removing entity-length limits — kalyan_kpl · 2026-08-25
- Unbounded Labs releases Bart, a vintage LLM trained on pre-1931 text — soggydoggy8 · 2026-08-25
- Mobile Benchmark: LFM2.5 Leads Efficiency on iPhone 17 Pro — ArtificialAnlys · 2026-08-25
- Distillation value: training Fable to distill a stronger Opus — HanchungLee · 2026-08-25