Harvey Legal Agent Benchmark Released: Top Models Achieve Only 14% Pass Rate
Artificial Analysis released the Harvey legal agent benchmark (LAB-AA), revealing that top models struggle with full task completion, achieving only up to 14.2% pass rates. Higher success strongly correlates with massive token generation, more agent loops, and longer processing times.
2026-07-08 ~ 2026-07-08 · 8 related posts
- Artificial Analysis Launches Harvey Legal Agent Benchmark — ArtificialAnlys · 2026-07-08
- Legal Tasks: Models Meet 90% of Criteria but Fail Full Pass — ArtificialAnlys · 2026-07-08
- Fable 5 Tops Legal Eval, GLM-5.2 Offers Best Value — ArtificialAnlys · 2026-07-08
- Legal Eval: Full Pass Rate Correlates with Token Output — ArtificialAnlys · 2026-07-08
- Stronger Models Take Longer on Agent Tasks — ArtificialAnlys · 2026-07-08
- Long-Range Agent Loops Become a Hallmark of Top Models — ArtificialAnlys · 2026-07-08
- Artificial Analysis Independently Reproduces Harvey Legal Eval — ArtificialAnlys · 2026-07-08
- Harvey Releases Legal Agent Benchmark Results — ArtificialAnlys · 2026-07-08