Grok 4.6 Tops GPQA Diamond Leaderboard with 94.9% Score
elonmusk · x · 2026-08-14
Elon Musk reshared a post highlighting Grok 4.6's latest benchmark performance. According to Artificial Analysis, Grok 4.6 secured the #1 spot on the GPQA Diamond benchmark with a score of 94.9%, beating out GPT-5.6, Gemini 3.1 Pro, and Claude Opus 5. Musk himself subsequently posted, encouraging users to 'Try Grok 4.6'.
Related event: xAI Launches Grok 4.6: Tops Benchmarks with Unmatched Cost-Efficiency(88 posts)→
More from Models
- Gemini 3.7 Flash Beats Grok 4.6 as Preferred Coding Model in Real-World Test — brandon_galang · 2026-08-14
- Gemini 3.7 Flash Tops Zapier's AutomationBench, Beating Pricier Claude and GPT Models — _philschmid · 2026-08-14
- Gemini 3.7 Flash Tested: Major Boosts in Coding and Reasoning at Half the Cost — petrusenko_max · 2026-08-14
- LiquidAI Models Hit Top 10 Trending on HuggingFace — JosephJacks_ · 2026-08-14
- Google's Gemini 3.7 Flash Pricing Strategy Baffles Developers — dbreunig · 2026-08-14
- Mistral Allegedly Shifts to Hosting Chinese LLMs, Unveils Moderation Model — SumitGup · 2026-08-14