NOHARM: an 1,100-task medical safety benchmark with an RCT of physician-AI teaming
davidjhwu · x · 2026-09-21
A Stanford-led team won Best Poster at the ANCO Genitourinary Cancers Symposium and released the preprint "First, do NOHARM": a medical safety benchmark paired with a randomized study.
- NOHARM (Numerous Options Harm Assessment for Risk in Medicine) is a 1,100-task benchmark of primary-care-to-specialist consultation cases
- It aims to quantify the frequency and severity of potentially harmful errors LLMs make in clinical consultations — an area where AI safety profiles remain poorly characterized
- The team also ran an RCT of physician-AI teaming on clinical consultations, with follow-up work building an AI safety oncology benchmark
A rare serious attempt to make "can LLMs harm patients" measurable via benchmark plus RCT, and an important reference for safe medical AI deployment.
Related event: Stanford's NOHARM Medical AI Safety Benchmark Wins Best Poster at ANCO(2 posts)→
More from AGI Musings
- "It's just vectors and multiplication": X users spar over whether understanding mechanisms rules out consciousness — cephaloform · 2026-09-21
- Fields Medalist Villani, June LLM skeptic, says OpenAI's math result left him shaken — zetalyrae · 2026-09-21
- Terence Tao: AI is cracking major math conjectures, "but we can afford to be slower" — rohanpaul_ai · 2026-09-21
- "AI writing is impossible to edit" — the real tell, plus 40-50 iteration reality — kfountou · 2026-09-21
- Shopify CEO Now Horrified by AI 'Slop Grenades' After Mandating AI Use — jonerp · 2026-09-21
- "Our language is not free data": rethinking ownership in language AI — danielleboyerr · 2026-09-21