Grok 4.6 hits 95% on GPQA Diamond, tied #1 and beating Opus 5 and GPT-5.6
XFreeze · x · 2026-08-28
xAI's Grok 4.6 (high) scores 95% on the GPQA Diamond leaderboard, tied for #1 with the highest score on the chart for graduate-level scientific reasoning. It outperforms Claude Opus 5, Fable 5 and GPT-5.6 Sol. GPQA Diamond tests models with science questions even experts find tough, putting Grok at the frontier of scientific reasoning.
More from Models
- GLM 5.3 Flash surges to 4th place in daily OpenRouter usage — Hesamation · 2026-08-28
- Tencent opens Hy4 preview weights: 770B MoE, 49B active, 1M context — Snoo26837 · 2026-08-28
- Similar benchmarks, double the size: Qwen3.8-Flash-Next needs 360GB vs DeepSeek-V4-Flash's 162GB lossless — vini542reddit · 2026-08-28
- Hands-on with Tencent Hy4 preview across 8 projects: better frontend taste, stable long-horizon tasks — vista8 · 2026-08-28
- PlayWorld Reveals Quality Gap in High-Scoring World Models — jiqizhixin · 2026-08-28
- Yoav Goldberg: home robots snapping to "standard" behavior fails everyday users too — yoavgo · 2026-08-28