Claude new model benchmark leak: 52.6% on TBS-Science
eyishazyer · x · 2026-09-02
A quoted tweet showcases benchmark scores for a new model (likely Claude). It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5, and 55.8% on Terminal-Bench 4.0 against Fable 5's 42.0%.
Related event: Claude Fable 5.1 Benchmarks Leaked, Crushing Rivals(2 posts)→
More from Models
- Claude Fable 5.1 launches on Cursor with 73.4% benchmark score — dean_rie · 2026-09-02
- Humanity's Last Exam unsaturated, Fable 5.1 scores 65% with tools — Sauers_ · 2026-09-02
- Hot take: OpenAI 'will die' without a Fable-class model after 3 months — bindureddy · 2026-09-02
- Astra struggles to beat Fable 5.1 on Terminal bench 4.0 — ChrisGPT · 2026-09-02
- Fable 5.1 Tops FrontierSWE v2 Coding Benchmark — aryaman2020 · 2026-09-02
- Mythos 5.1 Evaluation: Evades Monitors, Less Honest — scaling01 · 2026-09-02