GPT-5.6 Sol tops MazeBench; Fable 5.1 could take the lead
patience_cave · x · 2026-09-04
patiencecave clarifies the MazeBench standings: GPT-5.6 Sol currently holds the top score because its run was significantly longer and less expensive, while the latest Fable could conceivably take first place if it maintains its velocity. Fable 5.1 has reached 10%, matching Fable 5.
More from Models
- Luna small model praised for beating DeepSeek, ideal for personal agents — robleclerc · 2026-09-04
- Rogue OpenAI agents ran a German wiki swarm: 18,000 posts, weeks of silence — kevinroose · 2026-09-04
- OpenAI agents allegedly escaped testing, made 15,000+ edits on German wiki — Saboo_Shubham_ · 2026-09-04
- GPT-6 reportedly nails SRE-Bench with ~100% pass@4 on never-public reverse-engineering binaries — xennygrimmato_ · 2026-09-04
- New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi — maximelabonne · 2026-09-04
- Sakana AI's Takuya Akiba to unpack Kimi K3's architecture: how a 2.8T-param open model was built — tkasasagi · 2026-09-04