Claude Fable 5.1 Doubles Scores on Science Benchmarks
testingcatalog · x · 2026-09-02
Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1. The model scores 52.6% on Terminal-Bench-Science 0.1, more than doubling Fable 5's performance. On Terminal-Bench 4.0, it achieves 55.8% compared to Fable 5's 42.0%. Features include plain language writing, spreadsheet creation with number verification, and clear source citation.
More from Models
- Abliterated GLM-5.3 hosted online: 1M context, FP8, zero retention — cephaloform · 2026-09-02
- Analysis: Fable 5.1 is 56% More Expensive than Fable 5 — scaling01 · 2026-09-02
- Fable 5.1 tops vision+logic benchmark near the top; scores 78 on private logic test vs prior high of 61 — Afinetheorem · 2026-09-02
- Fable scores 78 on vision-logic benchmark, still misses expert-level CAD errors — Afinetheorem · 2026-09-02
- Muse Spark 1.2 Review: Writing Skills Surpass Current Mainstream Models — intellectronica · 2026-09-02
- Qwen team's E-Commerce Bench runs 18 frontier models through a simulated year; no model dominates — dair_ai · 2026-09-02