Study finds model rankings shift significantly on altered benchmark questions
abeirami · x · 2026-08-30
Research indicates that when benchmark questions are replaced with unreleased, slightly harder variants, some models exhibit significant drops in pass@3 scores, causing substantial shifts in model rankings. This raises questions about whether these models were 'benchmaxxed' and whether poor generalization beyond exact benchmarks is common. In contrast, other models maintain stable performance on the altered tests.
More from Models
- Qwen3.8-Flash-Next-REAP-288 Released in BF16 and GGUF Formats — EyalToledano · 2026-08-30
- GLM 5.3 Flash Inference Extremely Slow on Apple Silicon — CentrifugalMalaise · 2026-08-30
- OpenAI's unreleased Astra model leaks: first outputs of mozaik-alpha-fdm surface — gaganghotra_ · 2026-08-30
- Tencent open-sources Hunyuan Hy4: 770B MoE with 1M context — TencentHunyuan · 2026-08-30
- Tencent Hunyuan Hy4-preview runs in vLLM day 0: 770B MoE with 1M context — TencentHunyuan · 2026-08-30
- User complains about Anthropic's overly strict safety blocking on benign projects — BLUECOW009 · 2026-08-30