FrontierFinance eval: Fable 5.1 leads with 55.9% score, 1.7x cost increase
maithra_raghu · x · 2026-09-02
Partnered with Anthropic to evaluate Claude (Fable) 5.1 pre-launch on FrontierFinance. Fable 5.1 became the best performing frontier model with a score of 55.9%, significantly ahead of Fable 5 (49.2%). The gain largely comes from increased and better tool calls, such as anchoring on authoritative sources. However, this increased capability comes with a 1.7x cost increase compared to Fable 5.
More from coding & agent
- Optimizing 100K Token Skill Descriptions into MCP-Based ARD Server — dSebastien · 2026-09-02
- Perplexity uses Fable 5.1 as orchestrator with GPT 5.6 as cost-efficient subagents — AravSrinivas · 2026-09-02
- Prime Agent 0.9.1 ships with massive performance gains and many bug fixes — samsja19 · 2026-09-02
- Fable 5.1 Hands-On: Stronger Than Fable 5, Half the Tokens of Opus 5 — every · 2026-09-02
- Prathmesh Patel discusses agent reliability and testing — jeffiql · 2026-09-02
- Claude + Thrixel build full Roblox games from a single prompt — RanaHanocka · 2026-09-02