Clinical Triage Alignment Test: Closed Models Excel, Open-Source Lags
Researchers revisited a clinical triage alignment experiment after two years, utilizing the same set of 200 patient pairs to evaluate the consistency of frontier models from Google, OpenAI, and Anthropic in determining visit priorities.
已确认
Closed-source commercial models have achieved high levels of repeatability in clinical triage, with GPT-5.2 reaching 0.94, Claude Opus 4.8 at 0.92, Grok 4.5 at 0.89, and Gemini also scoring high. In contrast, open-weight models lag significantly behind: Llama 3.3 70B achieved an overall consistency (κ) of just 0.03 (0.13 for easy cases, and even negative for difficult ones), while GLM-4.7 scored 0.2. Furthermore, Grok 4.5 achieved a κ of 0.63 in an unaligned state, statistically comparable to the performance of Gemini and Opus.
为什么重要
This study confirms that the reliability and repeatability of closed-source models in rigorous tasks like medical triage are essentially resolved, whereas a substantial gap remains for open-source models. Additionally, in-context alignment technology significantly boosts the performance of frontier models on difficult problems (e.g., Grok 4.5 improved from 0.28 on hard pairs), while improvements on simpler cases are relatively limited.
2026-07-23 ~ 2026-07-23 · 6 related posts
Primary sources
- [source] A 200-pair triage test was rerun two years later on Google, OpenAI and Anthropic models — zakkohane · 2026-07-23
- [source] Closed models look highly reproducible, while open weights still lag badly — zakkohane · 2026-07-23
- [source] Open-weight models lag far behind on a triage test, with Llama 3.3 70B near zero — zakkohane · 2026-07-23
- In-context alignment mainly helps frontier models on the hard pairs — zakkohane · 2026-07-23
- Study finds Grok 4.5 matches Gemini and Opus on alignment-style tests — zakkohane · 2026-07-23
- Grok 4.5 tops a clinician-triage alignment test after in-context alignment — zakkohane · 2026-07-23