Grok 4.5 Leads Professional Task Benchmarks

elonmusk · x · 2026-07-10

New data from Snorkel reveals that Grok 4.5 outperforms other frontier models on GDPval+, a benchmark for real-world professional tasks. Covering expert-designed tasks across various economic sectors, Grok 4.5 achieved an average pass rate of 29%, surpassing GPT 5.5 at 22% and Claude Opus 4.8 at 21%. The post also highlights Grok 4.5's notable improvements in demanding fields like legal, education, healthcare, and QA analysis, showcasing xAI's focus on model performance in practical, actionable work.

Original post →

More from Models

Models channel →