LLMs hit gold-medal level on astronomy olympiad, Gemini 2.5 Pro scores 85.6%

bravo_abad · x · 2026-08-21

A new paper benchmarks five frontier LLMs on exams from the International Olympiad on Astronomy and Astrophysics — far harder than typical scientific QA, requiring deep conceptual understanding, multistep math derivations, and multimodal analysis, the capabilities that matter if LLMs are to act as scientific agents rather than retrieval systems.

Across four theory exams from 2022–2025, Gemini 2.5 Pro scored 85.6% and GPT-5 84.2%, both in gold-medal territory among 200–300 human participants; on data-analysis exams GPT-5 reached 88.5%, comparable to top-10 competitors. Yet the authors argue the error analysis is more informative than the scores: solving exams is still not doing science, and the models' mistakes show why.

Related event: GPT-5 and Gemini 2.5 Pro Win Gold at International Astronomy Olympiad Benchmark(2 posts)→

Original post →

More from Models

Models channel →