Grok Tops RareBench Diagnostics, DeepSeek v4-pro and GLM5.2 Underperform

ns123abc · x · 2026-08-14

Grok 4.6 has taken the crown on the RareBench benchmark for rare-disease diagnosis, edging out Anthropic's Claude Opus 5 at about 1/3 the cost. Meanwhile, DeepSeek's new v4-pro-0813 model underperformed expectations and its own v4-flash version, with potential API endpoint issues suspected. ZAI's GLM5.2 also underperformed, particularly in coding use cases compared to Kimi K3.

Related event: Grok 4.6 Tops Rare Disease Diagnosis Bench, DeepSeek Stumbles(2 posts)→

Original post →

More from Models

Models channel →