LLMs hit gold-medal level on astronomy olympiad, Gemini 2.5 Pro scores 85.6%
bravo_abad · x · 2026-08-21
A new paper benchmarks five frontier LLMs on exams from the International Olympiad on Astronomy and Astrophysics — far harder than typical scientific QA, requiring deep conceptual understanding, multistep math derivations, and multimodal analysis, the capabilities that matter if LLMs are to act as scientific agents rather than retrieval systems.
Across four theory exams from 2022–2025, Gemini 2.5 Pro scored 85.6% and GPT-5 84.2%, both in gold-medal territory among 200–300 human participants; on data-analysis exams GPT-5 reached 88.5%, comparable to top-10 competitors. Yet the authors argue the error analysis is more informative than the scores: solving exams is still not doing science, and the models' mistakes show why.
More from Models
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24
- Stealth Model 'Ox Alpha' Matches GPT-5.6 in Context Arena Tests — scaling01 · 2026-08-24
- Fable 5 Makes Up Only 6% of Anthropic Sales — rohanpaul_ai · 2026-08-24
- Opus 5 writes like Sorkin? How to fix wordy style — thk_ · 2026-08-24
- 2026 Chinese Model Landscape: DeepSeek V4, Kimi K3, and More Listed — TheTuringPost · 2026-08-24
- Rumor: Claude Opus 6 Leaked Specs Include 2.5M Context and Lower Costs — iamaliveix · 2026-08-24