o3 Tops the SciArena Leaderboard
allen_ai · x · 2026-07-16
According to an update from Allen AI, o3 ranks first on the SciArena leaderboard, outperforming Claude Opus 4.1, Gemini 3 Pro Preview, and open-weight models like DeepSeek-R1.
They noted that researchers found o3's responses to be more detailed and highly relevant. SciArena itself is an evaluation platform where researchers score answers to scientific literature questions.
Related event: Allen AI retires SciArena; o3 tops the science-QA leaderboard(6 posts)→
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22