Artificial Analysis Releases Harvey Legal Agent Benchmark Results

Artificial Analysis independently reproduced and released the full results of the Harvey Legal Agent Benchmark (LAB-AA), aiming to evaluate language models on real-world legal work across 24 practice areas. This benchmark reveals the capability boundaries and high computational costs of current top-tier models in complex legal tasks.

Evaluation Mechanism and Core Difficulties

The reproduced LAB-AA runs on Artificial Analysis's in-house Stirrup agent harness, supports context compression, and uses a simplified self-authored agent. The evaluation found that legal tasks demand high comprehensive capabilities: while leading models can meet over 90% of individual scoring criteria (only 4 models achieved this, including Fable 5, Opus 4.8, and GLM-5.2), they rarely satisfy all requirements for a complete task. The full pass rate ranges only from 0% to 14.2%.

Leaderboard and Cost-Effectiveness

Claude Fable 5 leads with a 14.2% full pass rate, but this top performance comes at a steep cost of about $18.9 per task, higher than Sonnet 5's approximate $11.8. In contrast, GLM-5.2 demonstrates the best cost-effectiveness. Furthermore, a model's full pass rate generally correlates with its workload per task: Fable 5 generated about 117,000 output tokens per task to achieve its pass rate, while Opus 4.8 and GLM-5.2 generated about 111,000 and 78,000 tokens, respectively.

Agentic Behavioral Analysis

Data indicates that long-range agentic loops have become a key feature of top models in complex tasks. Models leading in these tasks generally run more iterations: Fable 5 averaged about 64 loops per task, followed closely by Opus 4.8, GLM-5.2, and MiniMax-M3. Additionally, stronger models tend to spend more time per task; for instance, Fable 5 averaged about 16.9 minutes, Opus 4.8 about 18.5 minutes, and Sonnet 5 about 22.8 minutes.

2026-07-08 ~ 2026-07-08 · 8 related posts

Primary sources