ClinicalBench update shows Kimi K3 solving 7 of 10 EHR cases
teortaxesTex · x · 2026-07-23
A repost reacts to a ClinicalBench score update in which Kimi K3 solved 7 of 10 EHR sandbox cases, while Grok 4.5 remained ahead on the benchmark.
The quoted benchmark uses a virtual EHR sandbox where models must gather history, perform examination steps, order tests and procedures, and arrive at a diagnosis. The post highlights that one case stumped every model even though it was not especially hard for doctors.
More from Models
- Steve Hou expects a wave of U.S. open-source models as enterprise inference demand surges — soumitrashukla9 · 2026-07-23
- Musk says GPT-5 or GPT-6 could be indistinguishable from the smartest humans — kevinnbass · 2026-07-23
- One prompt was enough to get blocked, says an X user — gabriel1 · 2026-07-23
- Kimi K3 reportedly found and exploited a Redis 0day in 27 minutes with 32 agents — HanchungLee · 2026-07-23
- Reddit weighs a neglected MoE size class around 2B active parameters — WhoRoger · 2026-07-23
- Runescape Bench chart maps a crowded Pareto frontier across top models — teortaxesTex · 2026-07-23