Clinical Triage Retest: Closed-Source Models Far Outperform Open-Source
A researcher re-ran a clinical triage alignment experiment after two years, using the same 200 pairs of patient samples to evaluate the concordance of frontier models from Google, OpenAI, and Anthropic in determining triage priority.
Confirmed
Closed-source commercial models have reached a high level of repeatability in clinical triage, with GPT-5.2 achieving 0.94, Claude Opus 4.8 at 0.92, Grok 4.5 at 0.89, and Gemini also scoring high. In contrast, open-weight models lag significantly behind: Llama 3.3 70B's overall concordance (κ) is only 0.03 (0.13 for easy pairs, and negative for hard pairs), and GLM-4.7 is at 0.2. Additionally, Grok 4.5 achieved a κ of 0.93 in the unaligned state, which is statistically close to the performance of Gemini and Opus.
Why it matters
The study confirms that the reliability and repeatability of closed-source models for rigorous tasks like medical triage are essentially solved, but a massive gap remains for open-source models. Meanwhile, in-context alignment technology significantly boosts the performance of frontier models on hard problems (e.g., Grok 4.5 improving from 0.28 on hard pairs), while improvements on easy questions are relatively limited.
2026-07-23 ~ 2026-07-23 · 6 related posts
Primary sources
- [source] A 200-pair triage test was rerun two years later on Google, OpenAI and Anthropic models — zakkohane · 2026-07-23
- [source] Closed models look highly reproducible, while open weights still lag badly — zakkohane · 2026-07-23
- [source] Open-weight models lag far behind on a triage test, with Llama 3.3 70B near zero — zakkohane · 2026-07-23
- In-context alignment mainly helps frontier models on the hard pairs — zakkohane · 2026-07-23
- Study finds Grok 4.5 matches Gemini and Opus on alignment-style tests — zakkohane · 2026-07-23
- Grok 4.5 tops a clinician-triage alignment test after in-context alignment — zakkohane · 2026-07-23