Grok 4.5 Takes the Lead in Benchmark Scores
kevinnbass · x · 2026-07-10
The post states that Grok 4.5 significantly outperforms other frontier models, including Opus 4.8 and GPT 5.5, across multiple professional benchmarks, with particularly outstanding performance in legal tasks.
This serves as a direct capability comparison, highlighting its leading scores in a suite of professional benchmarks.
Related event: Grok 4.5 Tops Multiple Professional AI Benchmarks(5 posts)→
More from Models
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Leaker claims xAI is preparing Grok 4.7, hints at another surprise — mark_k · 2026-09-07
- Local LLMs now near Opus-level — what's still keeping them behind closed models? — mrsalvadordali · 2026-09-07
- Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77% — PawelHuryn · 2026-09-07
- Users miss the old Claude that used emojis: newer versions turn oddly poetic — JoshuaJosephson · 2026-09-07