FrontierFinance eval: Fable 5.1 leads with 55.9% score, 1.7x cost increase

maithra_raghu · x · 2026-09-02

Partnered with Anthropic to evaluate Claude (Fable) 5.1 pre-launch on FrontierFinance. Fable 5.1 became the best performing frontier model with a score of 55.9%, significantly ahead of Fable 5 (49.2%). The gain largely comes from increased and better tool calls, such as anchoring on authoritative sources. However, this increased capability comes with a 1.7x cost increase compared to Fable 5.

Original post →

More from coding & agent

coding & agent channel →