Fable 5.1 benchmarks double predecessor in coding and science tasks
sven_ai · x · 2026-09-02
Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5, and 55.8% on Terminal-Bench 4.0, up from 42.0%. These significant gains in coding and scientific reasoning capabilities indicate a rapid expansion in the scope of tasks AI Agents can handle.
Related event: Anthropic Releases Claude Fable 5.1 and Mythos 5.1(79 posts)→
More from Models
- Claude Fable 5.1 Preserved Thinking: Preventing Reuse of Reasoning Traces — illscience · 2026-09-02
- Test Shows Claude Fable 5.1 Finding Bugs Missed by Other Models — Sauers_ · 2026-09-02
- Fable 5.1 'pure insanity': user says it one-shotted his monthly usage limit — talkaboutdesign · 2026-09-02
- OpenAI's Astra rumored to use latent space reasoning, moving beyond text chains — Crazyscientist1024 · 2026-09-02
- Claude Fable 5.1 hits 90% on ARC-AGI-2; 275K-char system prompt leaked hours after launch — 新智元 · 2026-09-02
- OpenAI's 'recurrent depth' reasoning approach raises monitoring concerns — steph_palazzolo · 2026-09-02