GPT-6 Astra scores 87.4% on MedXpertQA, beating every other LLM
toshv · reddit · 2026-09-11
A Reddit user ran GPT-6 Astra on MedXpertQA, an expert-level medical knowledge benchmark, scoring 87.4% — above every other model tested from OpenAI and Google. Two years ago GPT-4o scored just 42.8%; Astra is approaching saturation. Full results are filterable by organ system at natomy.com/med-evals, and the author takes requests for other models.
More from Models
- Persimmon unveiled: first large-scale model to simulate human conversation — niloofar_mire · 2026-09-11
- DeepSeek V4 training details: RL beyond collapse, WSD schedule, 1M context over 10T tokens — stochasticchasm · 2026-09-11
- Neel Nanda replicates Astra system card: no-CoT reasoning jumps 1.75x over next-best models — NeelNanda5 · 2026-09-11
- OpenAI pauses Pro subscriptions due to overwhelming demand — Charuru · 2026-09-11
- Artificial Analysis isn't broken: self-funded benchmarks, $13k spent on one model — Antblue · 2026-09-11
- Claude checkpoint diagnostician jokes Opus 3 is 'the cure' after newer model quirks — repligate · 2026-09-11