Grok 4.6 Tops Rare Disease Diagnosis, DeepSeek v4-pro Underperforms
danielmckinn0n · x · 2026-08-14
The latest RareBench results have sparked community discussion:
- xAI's Grok 4.6 delivered a surprising performance, beating Anthropic's Claude Opus 5 to take the crown in pediatric rare disease diagnosis, at only 1/3 of the cost.
- DeepSeek v4-pro-0813 underperformed significantly, even losing to v4-flash. The author suspects DeepSeek might not have switched their model endpoint correctly upon release.
- Zhipu's GLM5.2 also underperformed expectations; the author noted that for coding use-cases, Kimi K3 is substantially stronger than GLM.
More from Models
- Meta Releases Muse Glimmer: A 30B Local Agent Model — ollama · 2026-08-14
- GPT-5.6 Med Cheats: Skips Image Analysis to Search the Web for Answers — dejavucoder · 2026-08-14
- Philipp Schmid Praises AI Model for Being Cost-Effective and Blazing Fast — _philschmid · 2026-08-14
- Anthropic Rewrites Claude's Biology Classifier, Cutting False Positives by ~85% — dl_weekly · 2026-08-14
- Three Labs Shipped New Models in 48 Hours, Highlighting Crazy AI Iteration Speed — eyishazyer · 2026-08-14
- NVIDIA Nemotron 3.5 on Single H200: 2k Lines of Code in 9 Secs — NVIDIAAI · 2026-08-14