Legal benchmark Pareto frontier: GPT-6 Luna at $0.22/task vs Claude at $18-22/task
ArtificialAnlys · x · 2026-10-09
Artificial Analysis published the score-vs-cost Pareto frontier for legal agent benchmark models with Hallucination-Gated All-Pass Rate above 0%: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh).
- Grok 4.7 (xhigh) leads at $9.50 per task
- Muse Spark 1.3 (max) second at $4.20
- The three Claude models cost $18 to $22 per task
- GPT-6 Luna (max) is cheapest at $0.22 per task, scoring 3.3%
More from Models
- Cohere Labs releases Tiny Aya L2-Thinker, a 3.35B model that reasons in 44 languages — lmoroney · 2026-10-09
- LightOnOCR-3 adds full-page grounding with one-prompt mode switching over OCR — IgorCarron · 2026-10-09
- TypeSafe AI's Jev decision model claims 200x faster, 400x cheaper classification in agent loops — LangChain · 2026-10-09
- Sentry CEO: Junior users report it 'seems smarter' after switching to Opus 5.5 — zeeg · 2026-10-09
- Polymarket opens GPT-6.1 Astra release market, 93% odds by year-end 2026 — Polymarket · 2026-10-09
- Zeeg says he's almost entirely stopped using GPT since Opus 5.5 launched — zeeg · 2026-10-09