Ground Truth Is Rarely Ground Truth in AI for Science, Researcher Warns

bravo_abad · x · 2026-09-15

Researcher bravoabad argues that in AI for Science, "ground truth" labels are rarely unquestionable truth: experimental labels carry instrument noise, calibration errors and systematic bias, while computational labels inherit the assumptions and failure modes of methods like DFT or molecular dynamics.

A DFT formation energy is not a measurement of nature — it's the output of a functional, and models trained on it inherit that functional's errors. A model can reproduce its training labels extremely well, look more precise than the measurements, and still be no more physically accurate: it learns the process that produced its targets, not nature itself.

His practical principle: before evaluating a model, ask what the target really is, how it was measured or computed, what uncertainty and approximations it carries, and whether it's the scientific quantity you actually care about — treat labels as scientific observations.

Original post →

More from AGI Musings

AGI Musings channel →