Science LLM benchmarks have flawed answers; fixing them significantly raises model scores
Profanion · reddit · 2026-09-16
A Reddit post highlights an arXiv paper (2609.13009) showing that many current science-focused LLM benchmarks contain errors in their reference answers. After correcting the flawed ground truths, model benchmark scores rose significantly.
This implies past leaderboard rankings on science tasks may have systematically underrated models, and that benchmark data quality is an overlooked variable when comparing models.
More from Models
- Essay claims Claude's personality converges on Anthropic's company culture — ryunuck · 2026-09-16
- ChatGPT Co-Inventor's New AI Startup Claims 200x Speed, Free Output Tokens Forever — Polymarket · 2026-09-16
- GPT-6 debugs Fallout 3's infamous Sentinel Lyons dialogue glitch — imjustnewatai · 2026-09-16
- That 1M-token context window can really burn your bill — peterjliu · 2026-09-16
- Speculation: top open-weight models may be distilling OpenAI and Anthropic, missing training code hints — dan_s_becker · 2026-09-16
- Aaronson hears AI companies have cracked longstanding TCS open problems, sitting on major announcements — scaling01 · 2026-09-16