NOHARM: an 1,100-task medical safety benchmark with an RCT of physician-AI teaming

davidjhwu · x · 2026-09-21

A Stanford-led team won Best Poster at the ANCO Genitourinary Cancers Symposium and released the preprint "First, do NOHARM": a medical safety benchmark paired with a randomized study.

A rare serious attempt to make "can LLMs harm patients" measurable via benchmark plus RCT, and an important reference for safe medical AI deployment.

Related event: Stanford's NOHARM Medical AI Safety Benchmark Wins Best Poster at ANCO(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →