Grok 4.6 tops MedAgentBench, edging out GPT-5.6 Sol on agentic clinical EHR tasks
elonmusk · x · 2026-08-18
Third-party evaluator MedicalSphereAI measured Grok 4.6 at 95.9% pass@1 (average of 3 runs) on MedAgentBench, a benchmark where the model acts as an autonomous agent in a simulated EHR, calling FHIR APIs across 10 clinical task types. That beats prior leader GPT-5.6 Sol (94.7%) and improves on Grok 4.5 by 2.5 points, with remarkably consistent runs (95.3%–96.3%). Musk shared the result.
More from Models
- User reports Qwen model overthinking despite low reasoning effort setting — Jebbyk1 · 2026-08-18
- User claims DeepSeek V4 Pro is mis-trained: cheaper Flash beats it — karminski3 · 2026-08-18
- Qwen3.8-27B Uncensored MLX build trends on Hugging Face for Apple Silicon — orcarouter · 2026-08-18
- Kimi K3 finds unpatched stack overflow in Go TS compiler — DanielLockyer · 2026-08-18
- Qwen3.8-27B Runs at 130 t/s Locally on Five-Year-Old Gaming GPUs — IgorCarron · 2026-08-18
- Sakana AI releases Japanese reasoning model Namazu on OpenRouter — SakanaAILabs · 2026-08-18