ATHENA Hits 94.7% Drug Reasoning Accuracy, Beating GPT-5 by 17.8 Points
marinkazitnik · x · 2026-07-07
On 3168 drug reasoning questions (DrugPC) and 456 patient-specific treatment cases (TreatmentPC), ATHENA achieved accuracies of 94.7% and 82.9% respectively, outperforming GPT-5 by 17.8 and 10.7 percentage points, and significantly surpassing DeepSeek-R1 (671B). The study found that providing GPT-5 with optional tool access barely improved its accuracy—it proactively used tools only about 1% of the time. This indicates that tool access itself is not the bottleneck; the key factor is whether the model can recognize its own knowledge boundaries and actively query information, rather than confidently providing answers when uncertain.
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11