Harvey's Legal Agent Benchmark shows conflicting scores: 19.6% vs 25.42%
3scorciav · x · 2026-10-01
A researcher noticed discrepancies in Harvey's Legal Agent Benchmark, which tests agents doing legal work with documents, spreadsheets and file-system tools: Harvey reports 19.6% for Argon, while vals.ai shows 25.42%, with Muse Spark 1.2 leading and Astra far behind. The post jokes that lawyers discussing cases over WhatsApp doesn't help benchmark clarity either.
More from Models
- Gemini 4 Argon unveiled: cyber-defense first, $2/$10 intro pricing — mrdbourke · 2026-10-01
- Law scholar: model spots unprompted Easter-egg joke, asks 'parrot camp' to explain — technollama · 2026-10-01
- Influencers hype Gemini 4 Argon: GOATED or hopelessly benchmaxed? — thatroblennon · 2026-10-01
- GPT 6 Series Shows Frequent Lazy Work, While Sol 5.6 Spent Hours on QA — jdjohnson · 2026-10-01
- Claim that DeepMind will beat Opus 5.5 at half price gets publicly called out as bogus — zacharynado · 2026-10-01
- Fulcrum's Echo claims to beat frontier models at style imitation with under $5K training cost — Hidenori8Tanaka · 2026-10-01