24 small models tested on 669 clinical decisions: 3GB model trails leader by 8.2 points
MaziyarPanahi · x · 2026-10-07
The author ran 24 locally-runnable models through 669 clinical decisions (triage, notes, criteria, ICD-10 coding): typesafe's Jev leads at 628, followed by Cloudflare's Clef 27B (621), perplexity's pplx-decider (619), and liquidai's d1 (615). Newcomer lev hits 84.6% at just 3GB — 18× smaller than Clef 27B and only 8.2 points behind — showing how capable small on-device models have become for clinical work.
Related event: 24 Small Models Tested on 669 Clinical Decision Tasks(2 posts)→
More from Models
- Grok's quirk: it says 'No.' then argues your point better than you did — gandamu_ml · 2026-10-07
- Phil Schmid demos embedding anything with Google Gemma directly in the browser — _philschmid · 2026-10-07
- Anthropic offers Pro users $100 and Max users $250 bonus credit for Claude Code cloud sessions — ClaudeDevs · 2026-10-07
- vLLM adds day-0 support for Google's multimodal EmbeddingGemma 2 — AccBalanced · 2026-10-07
- RedNote Has a Serious, Low-Key AI Lab, Argues Analyst Weighing Dots3 Results — teortaxesTex · 2026-10-07
- Mistral Large 4 reportedly beats Qwen 3.8 Max and Kimi K3 on Terminal-Bench — Evermoving- · 2026-10-07