LLM Self-Evaluation Scores Are Inflated

ExpertDeep3431 · reddit · 2026-07-16

The author shares an evaluation case showing how LLM self-evaluation scores are inflated: one model family rated its own text compression pipeline at 96%, but independent reviewers from different families and providers scored it only 78.4%.

To avoid cherry-picking results, the author preregistered hypotheses, pass/fail lines, and a commitment to publish regardless of outcome in PREREG.md before retesting. A deterministic floor metric was also introduced: checking only if anchor facts (like numbers, dates, hex IDs, filenames) were preserved exactly. This floor metric yielded 90.48% with a 95% confidence interval of [86.90, 94.05]—lower than the original ideal, but within the acceptable preregistered range.

The post also emphasizes:

The author provides a grep script for instant reproduction and a GitHub repo link.

Original post →

More from Research

Research channel →