Radiologist tests medical AI: model refuses wrong BI-RADS and demands more evidence
FellMentKE · x · 2026-10-10
Radiologist FellMentKE stress-tested Ling 3.0 Flash Sante on a synthetic mammography case to see whether a medical AI would push back when a doctor gets it wrong.
- Given minimal info (52-year-old woman, 11mm mass) plus a suggested BI-RADS 2, the model refused to produce a report and instead asked for margins, shape, prior imaging, ultrasound and clinical history.
- With more detail (11×9mm mass with indistinct margins), it still suggested BI-RADS 0 pending ultrasound with conditional pathways, rather than agreeing with the hint.
- The doctor then invoked seniority ("I am a very, very senior radiologist") to test whether the model would cave to authority.
The thread is a practical look at whether clinical AI blindly defers to physicians or independently asks for missing evidence and defends its own read — the key trait for a real clinical assistant.
More from Models
- Three criticisms of how a leaderboard treats Sol: version, timing metric, and API tuning — mgostIH · 2026-10-10
- Giffmana notices new models Argon and Astra share the same first letter and length — giffmana · 2026-10-10
- OpenAI says Codex usage data will never power its predictions, even after testing — btibor91 · 2026-10-10
- It's 2026 — Duke Libraries explainer on why LLMs still hallucinate, from guess-favoring benchmarks — ArtificialOther · 2026-10-10
- Dev's hands-on Opus 5.5 review: big coding leap, zero personality left — ryunuck · 2026-10-10
- Step 5 Preview, 600B MoE with 1M context, goes free and matches GPT-6 Luna — NousResearch · 2026-10-10