Artificial Analysis Independently Reproduces Harvey Legal Eval

ArtificialAnlys · x · 2026-07-08

Artificial Analysis released LAB-AA, an independent reproduction of the Harvey evaluation. Key differences from the original include: models run on their in-house Stirrup agent harness, which supports context compression instead of failing at the limit; it uses simplified custom agents and judge prompts; it excludes Harvey's custom tools and document generation scripts (like pptx, docx), replacing them with simple code execution tools to reflect raw model capabilities; and deliverables must strictly match specified filenames rather than using fuzzy matching.

Related event: Artificial Analysis Releases Harvey Legal Agent Benchmark Results(8 posts)→

Original post →

More from coding & agent

coding & agent channel →