Mazebench Launches to Stress-Test Fable 5.1 with Billions of Tokens
patience_cave · x · 2026-09-02
Following the Fable 5.1 release, developers have launched Mazebench, the world's toughest 3D spatial reasoning evaluation.
- Scale: A single run could take weeks and consume billions of tokens.
- Context: Fable 5 scored only 1% on this benchmark without Python.
The test aims to challenge the model's limits in high-dimensional spatial reasoning and long-horizon task handling.
More from Models
- Anthropic Fable 5.1 System Prompt Leaked, Spanning 270k+ Characters — Scobleizer · 2026-09-02
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Multi-agent evals lack model comparisons, need more details — scaling01 · 2026-09-02
- Fable 5.1 Beats GPT-5.6 on Benchmark at Lower Cost — haider1 · 2026-09-02
- Claude Fable 5.1 adds heavy instructions,疑似过度对齐 — teodorio · 2026-09-02
- Anthropic Cuts Cache Read Prices by 75%, Closing Gap with DeepSeek — Teknium · 2026-09-02