Grok 4.6 tops MedAgentBench, edging out GPT-5.6 Sol on agentic clinical EHR tasks

elonmusk · x · 2026-08-18

Third-party evaluator MedicalSphereAI measured Grok 4.6 at 95.9% pass@1 (average of 3 runs) on MedAgentBench, a benchmark where the model acts as an autonomous agent in a simulated EHR, calling FHIR APIs across 10 clinical task types. That beats prior leader GPT-5.6 Sol (94.7%) and improves on Grok 4.5 by 2.5 points, with remarkably consistent runs (95.3%–96.3%). Musk shared the result.

Original post →

More from Models

Models channel →