RadLE 2.0: A New Benchmark for Autonomous Medical Diagnosis

DrDatta_AIIMS · x · 2026-07-13

This thread introduces RadLE 2.0: a visual reasoning benchmark for autonomous diagnosis in radiology. It emphasizes that it's not just about "how many answers are correct," but whether the model correctly hands off to a human doctor when uncertain.

Several core metrics and leaderboards are provided:

The author's conclusion is clear: there is no single "best" model because the champion changes depending on the task objective. Furthermore, no model has yet reached the average human expert level across these metrics.

Related event: RadLE 2.0 Released: Benchmarking Medical AI Uncertainty(8 posts)→

Original post →

More from Models

Models channel →