MazeBench Results: Gemini 3.8 Flash Hits 4%, Most Models Under 1%
New MazeBench results show Gemini 3.8 Flash scoring 4% (up from 0% two months ago), most models below 1%, and GPT-5.6 Sol leading thanks to longer runtime and lower cost, with Fable potentially overtaking it.
2026-09-04 ~ 2026-09-04 · 4 related posts
- MazeBench results: Gemini 3.8 Flash scores 4%, most models under 1% in 3D open world — patience_cave · 2026-09-04
- MazeBench: Gemini 3.8 Flash scores 4%, Fable 5.1 matching its predecessor — patience_cave · 2026-09-04
- Gemini Flash agents improve world modeling: 0% to 4% in two months on MazeBench — patience_cave · 2026-09-04
- GPT-5.6 Sol tops MazeBench; Fable 5.1 could take the lead — patience_cave · 2026-09-04