Allen AI retires SciArena; o3 tops the science-QA leaderboard
On July 15, Allen AI retired SciArena, a benchmark that tested how well AI models answer questions about the scientific literature by having researchers judge the answers themselves. Beyond the final leaderboard, the project leaves behind a corpus of real scientific questions and expert-annotated ground truth intended to underpin more reliable science-QA evaluation. Around 1,700 users took part over its run.
Key results
Allen AI reported that o3 ranked first on the SciArena leaderboard, ahead of Claude Opus 4.1, Gemini 3 Pro Preview, and the open-weight DeepSeek-R1.
What researchers valued
SciArena's results show researchers weighted three factors most when judging answers: citation quality (23%), depth (19%), and whether the answer directly addressed the question (16%). Fluent prose alone did not win — in the science-literature setting, whether references are real, relevant, and verifiable mattered more.
Data assets and next steps
Allen AI framed SciArena's main value as collecting real scientific questions, expert preferences, and validated feedback into high-quality, expert-annotated ground truth. As AI agents are increasingly used to judge other AI agents, the team argues, evaluation itself must rest on reliable data. The analysis will be expanded in a NeurIPS 2025 Spotlight paper on how scientists evaluate AI-generated answers.
2026-07-16 ~ 2026-07-16 · 6 related posts
- [source] SciArena to Shut Down on July 15 — allen_ai · 2026-07-16
- Researchers Prioritize Citation Quality — allen_ai · 2026-07-16
- [source] o3 Tops the SciArena Leaderboard — allen_ai · 2026-07-16
- SciArena Offers Expert-Annotated Ground Truth — allen_ai · 2026-07-16
- [source] SciArena Paper: How Scientists Evaluate AI Answers — allen_ai · 2026-07-16
- SciArena Builds a Hub for Expert Evaluation Data — allen_ai · 2026-07-16