Study finds model rankings shift significantly on altered benchmark questions

abeirami · x · 2026-08-30

Research indicates that when benchmark questions are replaced with unreleased, slightly harder variants, some models exhibit significant drops in pass@3 scores, causing substantial shifts in model rankings. This raises questions about whether these models were 'benchmaxxed' and whether poor generalization beyond exact benchmarks is common. In contrast, other models maintain stable performance on the altered tests.

Original post →

More from Models

Models channel →