Grok 4.5 Leads Professional Task Benchmarks
elonmusk · x · 2026-07-10
New data from Snorkel reveals that Grok 4.5 outperforms other frontier models on GDPval+, a benchmark for real-world professional tasks. Covering expert-designed tasks across various economic sectors, Grok 4.5 achieved an average pass rate of 29%, surpassing GPT 5.5 at 22% and Claude Opus 4.8 at 21%. The post also highlights Grok 4.5's notable improvements in demanding fields like legal, education, healthcare, and QA analysis, showcasing xAI's focus on model performance in practical, actionable work.
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21