RadLE 2.0: A New Benchmark for Autonomous Medical Diagnosis
DrDatta_AIIMS · x · 2026-07-13
This thread introduces RadLE 2.0: a visual reasoning benchmark for autonomous diagnosis in radiology. It emphasizes that it's not just about "how many answers are correct," but whether the model correctly hands off to a human doctor when uncertain.
Several core metrics and leaderboards are provided:
- RadLE-C (Confidence-weighted): Claude Fable 5 leads, followed closely by Meta Muse Spark 1.1 and GPT-5.6 Sol Pro; the human expert baseline is higher.
- RadLE-A (Accuracy): Gemini 3.1 Pro takes first, GPT-5.6 Sol Pro second, and Meta Muse Spark 1.1 third.
- RadLE-R (Reliability): Claude Fable 5 approaches the human expert baseline.
- RadLE-S (Safety): Claude Fable 5 leads, but the overall field still lags behind humans.
- RadLE-H (Handoff Readiness): Meta Muse Spark 1.1 performs the best.
The author's conclusion is clear: there is no single "best" model because the champion changes depending on the task objective. Furthermore, no model has yet reached the average human expert level across these metrics.
Related event: RadLE 2.0 Released: Benchmarking Medical AI Uncertainty(8 posts)→
More from Models
- Frontier model weights are near-impossible to steal or self-replicate, argues Bindu Reddy vs AI doomers — bindureddy · 2026-09-11
- Why won't Google open source its STT models while startups ship SOTA open voice models? — techtotechbytechy · 2026-09-11
- User Praises DeepSeek's Model as Surprisingly Fast and Good in Hands-on Test — MaziyarPanahi · 2026-09-11
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11
- Pro 20x tier burns 60% of weekly quota in under a day with GPT-6 Astra — rschu · 2026-09-11
- Google isn't honoring its own Gemini Grounded Search pricing: only 289 of 15,000+ requests counted as free — ItalyExpat · 2026-09-11