ClinicalBench update shows Kimi K3 solving 7 of 10 EHR cases

teortaxesTex · x · 2026-07-23

A repost reacts to a ClinicalBench score update in which Kimi K3 solved 7 of 10 EHR sandbox cases, while Grok 4.5 remained ahead on the benchmark.

The quoted benchmark uses a virtual EHR sandbox where models must gather history, perform examination steps, order tests and procedures, and arrive at a diagnosis. The post highlights that one case stumped every model even though it was not especially hard for doctors.

Original post →

More from Models

Models channel →