Fable 5.1 doubles scientific benchmark score, surpassing Opus 5
felixrieseberg · x · 2026-09-02
On the Stanford-led Terminal-Bench-Science, Fable 5.1 scored 52.6%, more than doubling Fable 5 (24.7%) and beating Opus 5 (29.0%). This marks significant progress in AI scientific reasoning capabilities.
Related event: Fable 5.1 Doubles Science Benchmark Score and Refines Writing Style(3 posts)→
More from Models
- Fable 5.1 review: Tends to act as a 'manager' and plan globally — AlchainHust · 2026-09-02
- Elon Musk announces Grok 4.7 release in 10 days — XFreeze · 2026-09-02
- Report: OpenAI's Astra may use technique that destroys CoT monitorability — sjgadler · 2026-09-02
- Hands-on with Fable 5.1: taking over a 5+ day math problem across Codex and CC threads — DimitrisPapail · 2026-09-02
- Don't overreact negatively to Astra's recurrent depth — xuanalogue · 2026-09-02
- Deep dive into Astra's "recurrent depth" architecture trade-offs — daniel_mac8 · 2026-09-02