ATHENA Hits 94.7% Drug Reasoning Accuracy, Beating GPT-5 by 17.8 Points
marinkazitnik · x · 2026-07-07
On 3168 drug reasoning questions (DrugPC) and 456 patient-specific treatment cases (TreatmentPC), ATHENA achieved accuracies of 94.7% and 82.9% respectively, outperforming GPT-5 by 17.8 and 10.7 percentage points, and significantly surpassing DeepSeek-R1 (671B). The study found that providing GPT-5 with optional tool access barely improved its accuracy—it proactively used tools only about 1% of the time. This indicates that tool access itself is not the bottleneck; the key factor is whether the model can recognize its own knowledge boundaries and actively query information, rather than confidently providing answers when uncertain.
More from Research
- Krea 2 LoKr likeness guide says 750 steps is usually enough for near-perfect face training — LilBrownBebeShoes · 2026-07-22
- PoLar: Dynamically Skipping or Looping LLM Layers for Efficient Inference — ttkciar · 2026-07-22
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- ICML Tutorial: Is Optimization Theory Relevant in 2026? — srush_nlp · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22