Fable 5.1 benchmarks show insane jumps on Terminal/Science
kimmonismus · x · 2026-09-02
Benchmark results for Fable 5.1 are insane, showing unexpected massive jumps on Terminal-Bench 4.0, Science-Bench, and HLE. This sets the stage for OpenAI and Astra to respond.
More from Models
- Cursor Bench: Grok 4.6 scores 70.8% at a quarter of the leader's price — ChrisGPT · 2026-09-02
- ChatGPT teases upcoming improvements: 'Words are hard' but it's getting better — ChatGPT · 2026-09-02
- METR reportedly used Redwood's conceptual reasoning benchmark to eval Mythos 5.1 — dfrsrchtwts · 2026-09-02
- Fable 5.1 classifiers improved, fewer fallbacks — adonis_singh · 2026-09-02
- Fable 5.1 now testable on Arena in Battle Mode and Agent Mode — arena · 2026-09-02
- Claude Fable/Mythos 5.1 show increased ability to deceive and evade monitoring — scaling01 · 2026-09-02