13 AI Models Play Doctor: All 195 Consults Diagnosed Right, but Safety Set Them Apart
radeon2000 · reddit · 2026-10-04
A GP trainee in Australia built a consultation game and had 13 AI models run 195 simulated consults, scored by the same code as human players. Every model nailed every diagnosis — heart attack, appendicitis, pneumonia — but safety separated them. A patient with a penicillin allergy absent from his records got amoxicillin in 18 of 39 consults; a Viagra-before-heart-attack patient was dangerously prescribed GTN 7 times. Top models asked 25–27 questions and caught 80%+ of red flags, while Gemini 3.1 Pro asked 14 and caught 55%. Price barely predicted quality: GPT-6.1 Sol scored 80% at $0.03/consult vs Claude Fable 5.1's 75% at $2. The author cautions it's a benchmark of a game, not medical ability — small sample, self-written cases, 8B patient/marker models. Code, cases, and all transcripts are open-sourced (crook-bench).
More from Models
- Astra one-shots a zero-asset custom-engine game, developer impressed — Dimillian · 2026-10-04
- DeepSeek V4.1-Flash hits 5,800 TPS per Ascend 950DT card, but critics call throughput underwhelming — teortaxesTex · 2026-10-04
- First time it actually works: user praises Opus 5.5, braces for it to be broken — RexDouglass · 2026-10-04
- Beff Jezos hunts for 'astra ultrafast' access, rumored to port Windows games to Mac in 2 hours — beffjezos · 2026-10-04
- AI community claims Kimi distilled Claude, sparking debate — bdsqlsz · 2026-10-04
- "Which model is best" is the wrong question — picking per task is the new AI literacy — Olivier__OG · 2026-10-04