24 small models tested on 669 clinical decisions: 3GB model trails leader by 8.2 points

MaziyarPanahi · x · 2026-10-07

The author ran 24 locally-runnable models through 669 clinical decisions (triage, notes, criteria, ICD-10 coding): typesafe's Jev leads at 628, followed by Cloudflare's Clef 27B (621), perplexity's pplx-decider (619), and liquidai's d1 (615). Newcomer lev hits 84.6% at just 3GB — 18× smaller than Clef 27B and only 8.2 points behind — showing how capable small on-device models have become for clinical work.

Related event: 24 Small Models Tested on 669 Clinical Decision Tasks(2 posts)→

Original post →

More from Models

Models channel →