ATHENA Hits 94.7% Drug Reasoning Accuracy, Beating GPT-5 by 17.8 Points

marinkazitnik · x · 2026-07-07

On 3168 drug reasoning questions (DrugPC) and 456 patient-specific treatment cases (TreatmentPC), ATHENA achieved accuracies of 94.7% and 82.9% respectively, outperforming GPT-5 by 17.8 and 10.7 percentage points, and significantly surpassing DeepSeek-R1 (671B). The study found that providing GPT-5 with optional tool access barely improved its accuracy—it proactively used tools only about 1% of the time. This indicates that tool access itself is not the bottleneck; the key factor is whether the model can recognize its own knowledge boundaries and actively query information, rather than confidently providing answers when uncertain.

Related event: Open-Source Medical AI Agent ATHENA Outperforms GPT-5 in Treatment Reasoning(8 posts)→

Original post →

More from Research

Research channel →