Open-Source Medical AI Agent ATHENA Outperforms GPT-5 in Treatment Reasoning

Researchers introduced ATHENA, an open-source medical AI agent capable of treatment reasoning for all FDA-approved drugs since 1939. Instead of relying on static knowledge embedded in model weights, ATHENA dynamically deduces required evidence by invoking 212 biomedical tools, effectively addressing complex medical scenarios that require integrating disease backgrounds, comorbidities, and drug contraindications.

Performance and Benchmarks

On testing benchmarks, ATHENA achieved 94.7% and 82.9% accuracy on 3,168 DrugPC questions and 456 TreatmentPC cases, respectively. It outperformed GPT-5 by 17.8 and 10.7 points, and significantly surpassed DeepSeek-R1 (671B). Furthermore, in blind tests conducted across 28 rare disease expert organizations in the Biohub Rare As One network, ATHENA outperformed reference models across all 8 evaluation criteria, including cognitive traceability, covering cases like neurodevelopmental disorders and rare cancers.

Technical Mechanism and Training

ATHENA possesses a "metacognitive" ability to identify missing information and invoke the correct tools at each reasoning step. Due to the impossibility of manual annotation at this scale, the model was trained in two stages: the first involved automated generation of tool-calling trajectories for supervised fine-tuning, and the second utilized reinforcement learning with scientific feedback rewards. This enabled the model to learn to actively search for evidence before drawing conclusions, forming an inspectable reasoning chain.

Clinical Evaluation and Real-World Data Validation

In terms of clinical potential, ATHENA was evaluated by practicing physicians at Mount Sinai Hospital, successfully handling complex ICU-level cases lacking single guideline answers, such as post-bypass patients with chronic kidney disease. Additionally, the team validated ATHENA's adverse event association hypotheses using the longitudinal EHRs of 5.4 million patients from Israel's Clalit Health. The adjusted odds ratio for predicted associations ranged from 1.48 to 1.84, while negative controls remained near zero, proving the validity of its predictions in real-world data.

2026-07-07 ~ 2026-07-08 · 8 related posts