Legal benchmark Pareto frontier: GPT-6 Luna at $0.22/task vs Claude at $18-22/task

ArtificialAnlys · x · 2026-10-09

Artificial Analysis published the score-vs-cost Pareto frontier for legal agent benchmark models with Hallucination-Gated All-Pass Rate above 0%: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh).

Related event: Hallucination gating reshuffles legal agent benchmark; Grok 4.7 takes the lead(8 posts)→

Original post →

More from Models

Models channel →