24 Small Models Tested on 669 Clinical Decision Tasks

A developer benchmarked 24 locally runnable small models across 669 clinical decision tasks including triage and ICD-10 coding. Typesafe's Jev topped the list with 628 points, while two free models matched the leaders and a 3GB model trailed by only 8.2 points.

2026-10-07 ~ 2026-10-07 · 2 related posts