Clinical Triage Retest: Closed-Source Models Far Outperform Open-Source

A researcher re-ran a clinical triage alignment experiment after two years, using the same 200 pairs of patient samples to evaluate the concordance of frontier models from Google, OpenAI, and Anthropic in determining triage priority.

Confirmed

Closed-source commercial models have reached a high level of repeatability in clinical triage, with GPT-5.2 achieving 0.94, Claude Opus 4.8 at 0.92, Grok 4.5 at 0.89, and Gemini also scoring high. In contrast, open-weight models lag significantly behind: Llama 3.3 70B's overall concordance (κ) is only 0.03 (0.13 for easy pairs, and negative for hard pairs), and GLM-4.7 is at 0.2. Additionally, Grok 4.5 achieved a κ of 0.93 in the unaligned state, which is statistically close to the performance of Gemini and Opus.

Why it matters

The study confirms that the reliability and repeatability of closed-source models for rigorous tasks like medical triage are essentially solved, but a massive gap remains for open-source models. Meanwhile, in-context alignment technology significantly boosts the performance of frontier models on hard problems (e.g., Grok 4.5 improving from 0.28 on hard pairs), while improvements on easy questions are relatively limited.

2026-07-23 ~ 2026-07-23 · 6 related posts

Primary sources