Grok 4.5 and Muse 1.1 Top Legal Evaluation
alexandr_wang · x · 2026-07-10
Grok 4.5 and Muse Spark 1.1 achieved SOTA results on Harvey's LAB benchmark, demonstrating excellent performance in both cost and latency. Alexandr Wang recommended that legal professionals try out Muse Spark 1.1.
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21