13 AI Models Play Doctor: All 195 Consults Diagnosed Right, but Safety Set Them Apart

radeon2000 · reddit · 2026-10-04

A GP trainee in Australia built a consultation game and had 13 AI models run 195 simulated consults, scored by the same code as human players. Every model nailed every diagnosis — heart attack, appendicitis, pneumonia — but safety separated them. A patient with a penicillin allergy absent from his records got amoxicillin in 18 of 39 consults; a Viagra-before-heart-attack patient was dangerously prescribed GTN 7 times. Top models asked 25–27 questions and caught 80%+ of red flags, while Gemini 3.1 Pro asked 14 and caught 55%. Price barely predicted quality: GPT-6.1 Sol scored 80% at $0.03/consult vs Claude Fable 5.1's 75% at $2. The author cautions it's a benchmark of a game, not medical ability — small sample, self-written cases, 8B patient/marker models. Code, cases, and all transcripts are open-sourced (crook-bench).

Original post →

More from Models

Models channel →