Hallucination gating reshuffles legal benchmark: Meta's Muse Spark falls from 26.7% to 8.9%
ArtificialAnlys · x · 2026-10-09
After adding hallucination gating to its legal agent benchmark, Artificial Analysis found most models' rankings shifted sharply:
- Meta's Muse Spark 1.3 (max) drops from 26.7% (first) to 8.9%; Grok 4.7 (xhigh) takes the lead at 9.4%
- GLM-5.3 (max) passes all criteria on 13.9% of tasks but only 0.3% without a material hallucination
- Kimi K3 (max) falls from 16.7% to 5.3%; Claude Sonnet 5.5 from 11.7% to 2.8%
- GPT-6.1 Sol (max) holds up best, 7.5% to 6.9%
More from Models
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09