Gemini 2.5 Pro hits 85.6% on IOAA theory; geometry still costs all models 15-26 points

2026-08-21

Five models sat 57 IOAA problems (2022-2025). Gemini 2.5 Pro scored 85.6% on theory, GPT-5 84.2%, both gold; only GPT-5 held 88.5% on data analysis. Geometry cost 15-26 points.

What problem this solves

Astronomy already has LLM demos for gravitational-wave searches and multi-band galaxy interpretation. Those pipelines show that a task can be wired up. They do not measure whether a model can derive, approximate, read a plot, or reason about the celestial sphere. Existing quizzes such as AstroBench and Astro-QA mostly test recall with multiple choice.

A team at Ohio State and the University of São Paulo turned the International Olympiad on Astronomy and Astrophysics (IOAA) into a benchmark. IOAA draws about 200-300 high-school contestants a year. The syllabus covers cosmology, spherical trigonometry, stellar astrophysics, celestial mechanics, photometry, and instrumentation. The observation round needs telescopes and star charts, so it was dropped.

Method

The set is 49 theory problems and 8 data-analysis problems from 2022-2025. Models saw the original LaTeX, embedded figures, and the same constants sheet given to students: 16 physical constants, 26 astronomical quantities, and 6 calculus identities. The prompt demanded step-by-step reasoning, LaTeX math, and tikz/pgfplots figures. Incomplete reasoning scored zero. One attempt per problem, no retries.

Two IOAA officials graded against the official rubrics. Lucas Carrit Delgado Pinheiro was a 2018 contestant, team leader in 2022, 2023 and 2025, and academic-committee member in 2024; Bruno Caixeta Piazza has a similar record. They scored 285 scripts (57 problems × 5 models) independently and resolved every disagreement. Accommodations covered the medium, not the physics: a failed tikz plot still scored if the caption described the right figure; LaTeX syntax noise that did not change the math was ignored. A numerical slip was deducted once, not cascaded.

The five models are GPT-5, Gemini 2.5 Pro, OpenAI o3, Claude Opus 4.1, and Claude Sonnet 4. The latest knowledge cutoff is March 2025. IOAA 2025 was held in August, so that year is clean. Scores on 2025 sat close to each model's four-year mean, which the authors read as limited contamination on 2022-2024.

Theory items split into Category I (celestial geometry, spherical trig, coordinate transforms; about 37% of items) and Category II (astrophysical calculation without spatial visualization; 63%). Difficulty follows the human median: Easy (median above 50%) 21%, Medium 25%, Hard 35%, Extra Hard 19%. Medals follow IOAA's relative rule: bronze at 100-130% of the median, silver 130-160%, gold above 160%. Because observation was omitted, gold cutoffs were computed separately on theory and on data analysis.

Results

On theory, Gemini 2.5 Pro averages 85.6% (±8.0) and GPT-5 84.2% (±6.1). The pack behind them is thinner: o3 77.5%, Opus 64.7%, Sonnet 60.6%. GPT-5 led in 2022 (93.0%), 2023 (89.6%), and 2025 (86.8%). 2024 put 63% of points in geometry; Gemini took 83.0% and GPT-5 fell to 67.5%.

The human bar is harsher than the raw percentages. The 2025 theory gold line was 31.3%. All five models cleared gold. GPT-5's 86.8% ranked first, 443% of the median. GPT-5 beat the best human on theory in 2022, 2023, and 2025; Gemini did the same in 2022 and 2023. The only silver was Claude Sonnet 4 in 2023 (57.4%, rank 62).

Data analysis splits the field. GPT-5 averages 88.5%, above its own theory score, gold every year and inside the top 10 each time; it beat the best student in 2022 and 2023. Everyone else dropped 10-15 points from theory: Gemini 75.7%, o3 67.7%, Opus 54.8%, Sonnet 47.9%. In 2024 and 2025 the Claude models landed at bronze or no medal. Sonnet scored 30.0% on 2025 analysis, 84.4% of the median, rank 177.

The hole is geometric. Category II physics sits at 89-91% for the top three. Category I geometry sits at 78.6% (Gemini), 76.1% (GPT-5), and about 52% for both Claudes, a 15-26 point gap. Conceptual errors plus geometric/spatial errors account for 60-70% of theory points lost. Arithmetic, notation, and approximation cost little.

The misses are specific. On 2024 theory Q10 (eclipse geometry), only GPT-5 and Gemini treated the Sun, Moon, and Earth as generally non-collinear; the other three assumed collinearity and the whole solution died. On 2025 Q1, given a 120° detector angle, the answer is 30°; every model except Gemini returned 60°. On 2025 Q2, no model picked tropical versus sidereal year correctly in the first subpart. Models can recite Wien's law and still pick the wrong blackbody curve from a figure.

Why it matters

For people building astronomy tools, this replaces "does the model know astronomy" with a human-calibrated olympiad. Use the result in pieces. Formula checks, parameter scans, and concept cross-checks are already in range for the top models. Independent plot reading, spherical geometry, and deciding when an approximation is legal are not.

The paper sits with the IMO, IPhO, and IOI evaluations, adding astronomy. On AstroBench these models look closer together. On IOAA derivations, Opus trails GPT-5 by nearly 20 points. A knowledge quiz hides a reasoning crack that this exam opens.

This is a careful grading study, not a new method. The contribution is the human bar: official mark schemes, two real olympiad examiners, and a 2025 contamination check.

Limitations

The observation round was skipped. Real IOAA gold uses the sum of three papers against the median. Gold here is a per-paper cutoff, not a live three-paper total against the field. Rankings insert each model into the student list separately. Each problem was run once, so there is no sampling variance. Eight data-analysis items is a thin base for multimodal claims. 2022 Category I scores look too high: only 4 of 13 theory items were geometric, and one long problem restates a 1981 paper on tadpole and horseshoe orbits. Both graders are co-authors.

An olympiad, even a hard one, still has a mark scheme. Gold on a written paper is not an autonomous research agent.

Terms

Source

What people are saying

All paper explainers