Science LLM benchmarks have flawed answers; fixing them significantly raises model scores

Profanion · reddit · 2026-09-16

A Reddit post highlights an arXiv paper (2609.13009) showing that many current science-focused LLM benchmarks contain errors in their reference answers. After correcting the flawed ground truths, model benchmark scores rose significantly.

This implies past leaderboard rankings on science tasks may have systematically underrated models, and that benchmark data quality is an overlooked variable when comparing models.

Original post →

More from Models

Models channel →