RadLE 2.0: A Benchmark for Autonomous Medical Diagnosis

alexandr_wang · x · 2026-07-14

The post shares the release of RadLE 2.0, a visual reasoning benchmark for autonomous medical diagnosis that emphasizes uncertainty awareness and the ability to "know when to stop and hand off to a human."

It also notes recent advancements in frontier models from OpenAI, Meta, and xAI. The team has evaluated frontier, open-source, and medical VLMs on RadLE 2.0 and published a leaderboard.

The core question the author wants to highlight isn't simply "who scores highest," but rather: before allowing AI to take full autonomous control over diagnostics, does the model know when it should stop and defer to a human?

Related event: RadLE 2.0 Released: Benchmarking Medical AI Uncertainty(8 posts)→

Original post →

More from Models

Models channel →