Fable 5.1 beats Fable 5, matches Opus 5 on ML bench as refusals drop to 0/12
xeophon · x · 2026-09-02
User wassname ran their own wassname-ml-bench to test whether Fable 5.1 was nerfed for ML: it outperforms Fable 5 and is on par with Opus 5 at machine learning. Anthropic also improved refusal classifiers — Fable 5 went from 2/12 to 0/12 refusals, Fable 5.1 sits at 1/12. The test responds to a question citing Anthropic's note on improved cyber/bio classifiers for Fable 5.1, and whether AI-research classifiers still throttle capability.
More from Models
- Gary Marcus clashes with OpenAI researcher over CoT monitoring and unmonitorability race — GaryMarcus · 2026-09-02
- Tencent open-sources WeMM-Embedding: unified text/image/video embeddings, Apache 2.0 — tomaarsen · 2026-09-02
- When nobody can track frontier model progress, closed-model business may lose to open weights — StewartalsopIII · 2026-09-02
- Looped transformer is no dark art: rasbt debunks the OpenAI Astra rumor — rasbt · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Anthropic's Fable-5.1-max Grabs #4 on eyebench-v3 in Biggest Bench Jump Yet — adonis_singh · 2026-09-02