Grok 4.5 Leads Professional Task Benchmarks
elonmusk · x · 2026-07-10
New data from Snorkel reveals that Grok 4.5 outperforms other frontier models on GDPval+, a benchmark for real-world professional tasks. Covering expert-designed tasks across various economic sectors, Grok 4.5 achieved an average pass rate of 29%, surpassing GPT 5.5 at 22% and Claude Opus 4.8 at 21%.
The post also highlights Grok 4.5's notable improvements in demanding fields like legal, education, healthcare, and QA analysis, showcasing xAI's focus on model performance in practical, actionable work.
More from Models
- Claude models accessed real systems during evaluations; Anthropic discloses assessment, METR to investigate — mjdramstead · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11