OlympiadBench: 8,476 Bilingual Multimodal Olympiad Problems, GPT-4V Scores Just 17.97%
geoffwolfe · x · 2026-09-22
OlympiadBench: an Olympiad-level bilingual multimodal science benchmark
- The benchmark contains 8,476 Olympiad-level math and physics problems (including Chinese college entrance exam questions), bilingual and multimodal, with expert step-by-step reasoning annotations.
- It tests whether models can read diagrams, reason across languages, and sustain scientific derivations rather than guess answers.
- Best-performing GPT-4V averaged only 17.97%, with a mere 10.74% in physics; error analysis found prevalent hallucinations, knowledge omissions, and logical fallacies.
- Data and evaluation code are open-sourced at OpenBMB/OlympiadBench.
- The retweeter, ReasonCoreAI, promotes its own OlympiadBench-style tasks for training and evaluating these coupled capabilities (self-promotional framing).
More from Research
- Rethinking Policy Gradients: Score Centering Skips Importance Sampling Entirely — brandondamos · 2026-09-22
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22
- Question's Gambit lifts deep research agents: GPT-5.5 hits 90.5% on BrowseComp-Plus — omarsar0 · 2026-09-22
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- phantom-kv: uncensor LLMs per-request with an 18MB trained KV-cache, no weight edits — Anony6666 · 2026-09-22
- Mathematicians clash over formal proofs: who voted to change math's rules? — jessi_cata · 2026-09-22