MazeBench Update: Grok explores most, Gemini Flash solves puzzles while others struggle with T-shaped blocks
patience_cave · x · 2026-08-25
New MazeBench results for coding agents:
- Grok 4.6: Explored the most rooms (6), but scored 0% on solving puzzles.
- Gemini 3.7 Flash: Scored 1%, the only model to solve puzzles (collecting gems), outperforming Kimi K3.
- Others: ox-alpha, glm, and qwen all scored 0%.
Behavioral Observations:
- ox-alpha, glm, and qwen struggled with a "T-shaped" block, failing to push it down a correspondingly shaped hole.
Related event: MazeBench: Most Flagship Models Score Zero on Maze Tasks(3 posts)→
More from coding & agent
- Introducing Wake: A Rust-based Multiplayer AI Coworker OS for Teams and Agents — prasannaalahoti · 2026-08-25
- Free CLI tool launches to scan LLM endpoints for 15 prompt-injection attacks — Ventrovadev · 2026-08-25
- Case study: Two LLMs missed a future-data bug in coding and review loop — niacolhealth · 2026-08-25
- Jeffrey's Skills launches CLI tool for premium AI coding workflows — doodlestein · 2026-08-25
- Full Workflow for Optimizing Rust Code with 0x Alpha Model — doodlestein · 2026-08-25
- Build a voice agent with LangGraph and ElevenLabs — dl_weekly · 2026-08-25