GPT-5 and Gemini 2.5 Pro Win Gold at International Astronomy Olympiad Benchmark

A Nature Astronomy study benchmarked five frontier LLMs on IOAA problems requiring deep reasoning, multi-step math, and multimodal analysis, finding gold-medal-level performance, with Gemini 2.5 Pro scoring 85.6%.

2026-08-20 ~ 2026-08-21 · 2 related posts

1 near-duplicate retellings: bravo_abad