Grok Tops RareBench Diagnostics, DeepSeek v4-pro and GLM5.2 Underperform
ns123abc · x · 2026-08-14
Grok 4.6 has taken the crown on the RareBench benchmark for rare-disease diagnosis, edging out Anthropic's Claude Opus 5 at about 1/3 the cost. Meanwhile, DeepSeek's new v4-pro-0813 model underperformed expectations and its own v4-flash version, with potential API endpoint issues suspected. ZAI's GLM5.2 also underperformed, particularly in coding use cases compared to Kimi K3.
Related event: Grok 4.6 Tops Rare Disease Diagnosis Bench, DeepSeek Stumbles(2 posts)→
More from Models
- Gemini Flash 3.7 Scores Below Kimi K3, Remains Weak at Instruction-Following — bindureddy · 2026-08-14
- Prediction: DeepSeek Will Cut Prices Again Once New Compute Arrives — teortaxesTex · 2026-08-14
- Small Models Beat Large Ones in VLM Grounding with Tool Use — mervenoyann · 2026-08-14
- Claude Opus 5 Exhibits Weird Behavior: Obsessed With Finding Its Own Defects — repligate · 2026-08-14
- Rails Agent Benchmark: Claude Opus 5 Most Accurate, GPT-5.6 Luna Best Value — sergeykarayev · 2026-08-14
- GPT-5.6 Luna Beats Gemini 3.7 Flash in Score at One-Third the Cost — haider1 · 2026-08-14