ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without
geoffwolfe · x · 2026-09-27
Researchers from Qiyuan Tech, Peking University, Tsinghua and HKU introduced ScienceArena, a benchmark built on 13 years of recent science olympiad contests to counter benchmark saturation and data contamination. Headline finding: chemistry turns requiring structural diagrams earn only 64.5% of available points versus 74.1% without them — models can narrate plausible reaction stories but falter when they must commit to a precise molecular structure. The authors argue evals should grade actual structural commitment, not fluent scientific prose.
More from Models
- Yacine Still Can't Stand Opus 5.5's Writing Style — yacineMTB · 2026-09-27
- Google's AI lineup looks stalled: nano banana untouched since June, Gemini Pro since February — haider1 · 2026-09-27
- Rumor: partners got a better Claude Sonnet 5 checkpoint, release expected next week — kimmonismus · 2026-09-27
- Codex surprise hard reset angers users: 60% saved quota wiped, reset pushed 7 days out — ChrisUniverse · 2026-09-27
- Ex-NVIDIA engineer: US labs ignored global users, so the world runs on Chinese models — ivan_bezdomny · 2026-09-27
- GPT-OSS chat template bug silently drops past answers, degrading multi-turn coherence — arbv · 2026-09-27