Red flag: a new model crushing old benchmarks but not new ones likely signals contamination
rajammanabrolu · x · 2026-10-06
rajammanabrolu shares a heuristic for evaluating new models: be wary when a model is much better on significantly older benchmarks in a category but not on newer ones — it raises the odds of confounders like training data contamination or benchmark overfitting.
More from Models
- Tester claims ChatGPT 6 Astra (medium) burns far more usage than Ultra tier — ChrisUniverse · 2026-10-06
- ChatGPT has 3x Claude's paid subscribers as Claude's US growth cools — FinanceYF5 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06
- StealthGPT Launches Super 'Humanizer' Model, Claims It Beat AI Detector Pangram — menhguin · 2026-10-06
- Reflection Ships Apache 2.0 Model With Tech Report; Analyst Estimates Pre-training MFU at Just ~12% — eliebakouch · 2026-10-06
- New paper probes Olmo 3 puzzle: near-identical short-context models diverge after long-context extension — kylelostat · 2026-10-06