Harvey's Legal Agent Benchmark shows conflicting scores: 19.6% vs 25.42%

3scorciav · x · 2026-10-01

A researcher noticed discrepancies in Harvey's Legal Agent Benchmark, which tests agents doing legal work with documents, spreadsheets and file-system tools: Harvey reports 19.6% for Argon, while vals.ai shows 25.42%, with Muse Spark 1.2 leading and Astra far behind. The post jokes that lawyers discussing cases over WhatsApp doesn't help benchmark clarity either.

Original post →

More from Models

Models channel →