RadLE 2.0: A New Benchmark for Autonomous Medical Diagnosis
DrDatta_AIIMS · x · 2026-07-13
This thread introduces RadLE 2.0: a visual reasoning benchmark for autonomous diagnosis in radiology. It emphasizes that it's not just about "how many answers are correct," but whether the model correctly hands off to a human doctor when uncertain.
Several core metrics and leaderboards are provided:
- RadLE-C (Confidence-weighted): Claude Fable 5 leads, followed closely by Meta Muse Spark 1.1 and GPT-5.6 Sol Pro; the human expert baseline is higher.
- RadLE-A (Accuracy): Gemini 3.1 Pro takes first, GPT-5.6 Sol Pro second, and Meta Muse Spark 1.1 third.
- RadLE-R (Reliability): Claude Fable 5 approaches the human expert baseline.
- RadLE-S (Safety): Claude Fable 5 leads, but the overall field still lags behind humans.
- RadLE-H (Handoff Readiness): Meta Muse Spark 1.1 performs the best.
The author's conclusion is clear: there is no single "best" model because the champion changes depending on the task objective. Furthermore, no model has yet reached the average human expert level across these metrics.
Related event: RadLE 2.0 Released: Benchmarking Medical AI Uncertainty(8 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22