Fable 5 Tops Legal Eval, GLM-5.2 Offers Best Value
ArtificialAnlys · x · 2026-07-08
Artificial Analysis's Harvey legal evaluation shows that topping the leaderboard is expensive. Claude Fable 5 leads with a 14.2% full pass rate but is also the most expensive model at $18.9 per task, higher than Sonnet 5's $11.8; Opus 4.8 costs $8.2 per task. In terms of cost-effectiveness, DeepSeek V4 Flash (max) passed 1.7% of tasks for only $0.08, while GLM-5.2 (max) matched Opus 4.8 at roughly 15% of the cost ($1.3 per task).
Related event: Artificial Analysis Releases Harvey Legal Agent Benchmark Results(8 posts)→
More from Models
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22