24 Small Models Tested on 669 Clinical Decision Tasks
A developer benchmarked 24 locally runnable small models across 669 clinical decision tasks including triage and ICD-10 coding. Typesafe's Jev topped the list with 628 points, while two free models matched the leaders and a 3GB model trailed by only 8.2 points.
2026-10-07 ~ 2026-10-07 · 2 related posts
- 24 models tested on 669 clinical decisions: Jev stays #1 as two free models close in — MaziyarPanahi · 2026-10-07
- 24 small models tested on 669 clinical decisions: 3GB model trails leader by 8.2 points — MaziyarPanahi · 2026-10-07