Claude Fable 5 scores 210/210 on a bar exam benchmark for about $6
ctjlewis · x · 2026-07-22
Claude Fable 5 hits 210/210 on a bar exam benchmark
Matthew Stubenberg says Anthropic’s Claude Fable 5 became the first model to score perfectly on 210 MBE bar exam practice questions in his benchmark.
- The run cost about $6 in API tokens.
- It averaged roughly 10 seconds per question.
- Stubenberg notes the benchmark started in November 2022, when ChatGPT 3.5 was around 50% on the same set.
- His chart shows a steady climb toward perfect scores, with Fable 5 reaching 100%.
He frames it as evidence of how fast model performance on professional exams has improved in just over three years.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11