pplx-decider hits 95.2% on 669 clinical decisions, ties Jev at #1
MaziyarPanahi · x · 2026-10-07
MaziyarPanahi updated his clinical decision benchmark: 24 Jev-style models evaluated on 669 clinical decisions covering triage, notes, criteria, and ICD-10 coding.
- The new pplx-decider (Perplexity) scored 637/669 (95.2%), the top score, tied #1 with Jev
- Jev (@typesafeai) previously led with 628; two free models caught up
- Second tier: Clef 27B (Cloudflare) at 621, d1 (Liquid AI) at 615
- He's taking suggestions on which model to run next
More from Models
- Developer slams Claude for refusing piano sheet music, calls out Anthropic's copyright double standard — tetsuoai · 2026-10-07
- Leaker predicts Anthropic will ship something within days to steal OpenAI's spotlight — Dr_Singularity · 2026-10-07
- Claude subscription rumored to add Midjourney and Suno, with Opus 5.5 powering tasks — op7418 · 2026-10-07
- Perplexity's open-licensed pplx-decider-v1-27b, a Qwen3.8-27B finetune, trends on Hugging Face — perplexity-ai · 2026-10-07
- Temp 0 doesn't guarantee deterministic LLM output — batch shape and kernels shift logits — JFPuget · 2026-10-07
- OpenAI doesn't need more resets, needs better communication, argues paying user — AirportEither2456 · 2026-10-07