24 models tested on 669 clinical decisions: Jev stays #1 as two free models close in
MaziyarPanahi · x · 2026-10-07
The author benchmarked 24 models across 669 clinical decision tasks — triage, clinical notes, criteria checks, and ICD-10 coding. Jev retained the top spot with 628 points, but Cloudflare's Clef 27B (621) and Perplexity's pplx-decider (619), both free, nearly closed the gap; Liquid AI's d1 scored 615. A sign that free models are now competitive with paid leaders on healthcare-adjacent tasks.
Related event: 24 Small Models Tested on 669 Clinical Decision Tasks(2 posts)→
More from Models
- Mistral Large 4 'Le Chonk': 1T params, 49B active, Europe's best open-weights model — giffmana · 2026-10-07
- Reflection Beam and Mistral Large 4 hit GLM-5.2 level, sparking distillation gap debate — Yuchenj_UW · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- Decider model gains from unmasking and more data, authors deny benchmaxxing — antoine_chaffin · 2026-10-07
- Small public-private leaderboard gap cited as proof of no benchmaxxing — antoine_chaffin · 2026-10-07
- Tesla's Grok voice goes hoarse but can't hear itself — a look at AI engineering shortcuts — PTrubey · 2026-10-07