Medical Diagnostic Benchmark RadLE 2.0 Released

shuyanzh36 · x · 2026-07-14

This content focuses on Radiology’s Last Exam 2.0 (RadLE 2.0), a visual reasoning benchmark for autonomous AI diagnostics in radiology that incorporates uncertainty awareness.

Key highlights include:

Reposts noted that Muse Spark 1.1 outperformed GPT-5.6 Sol and Gemini 3.1 on RadLE, though it still trails Fable. Human doctors currently remain superior. The overarching argument: before granting AI higher autonomy, self-awareness regarding limitations is more critical than raw scores.

Related event: RadLE 2.0 Released: Benchmarking Medical AI Uncertainty(8 posts)→

Original post →

More from Research

Research channel →