What Does Sante's 83.83 on DiagnosisArena-MCQ Actually Measure?
Expert_Coffee_203 · reddit · 2026-09-09
A Reddit thread questions what Sante's 83.83 score on DiagnosisArena-MCQ really indicates: the benchmark aggregates several task types, so the score may not translate directly to general medical reasoning ability. The poster asks how much weight such scores deserve when comparing medical reasoning models.
More from Research
- Dev Spotlight: Transformers Forced to Predict Themselves and Build Belief States — yacinelearning · 2026-09-09
- Open-source 4B VLM RL-trained to play GeoGuessr beats GPT 5.4 mini and Haiku on one A100 — SergioPaniego · 2026-09-09
- Granularity and the 'birdiest bird': self-supervised clusters cut across ImageNet labels — y_m_asano · 2026-09-09
- Reviewer finds AI-written peer reviews rampant at ARR, calls for benchmarks of AI reviews — ChenhaoTan · 2026-09-09
- MirroS finds S-Space: a manipulable spatial workspace inside multimodal models — HeyAmit_ · 2026-09-09
- Synthetic Data May Leak Answers, Security Expert Questions — matthew_d_green · 2026-09-09