Red flag: a new model crushing old benchmarks but not new ones likely signals contamination

rajammanabrolu · x · 2026-10-06

rajammanabrolu shares a heuristic for evaluating new models: be wary when a model is much better on significantly older benchmarks in a category but not on newer ones — it raises the odds of confounders like training data contamination or benchmark overfitting.

Original post →

More from Models

Models channel →