IBM Study: Benchmark High Scores Depend on Wording, Strong Models More Fragile
omarsar0 · x · 2026-08-16
New IBM research reveals that model performance on benchmarks is heavily dependent on question phrasing. Using BenchDrift to generate meaning-preserving variations of problems, the study finds that weaker models often benefit from rephrasing, while stronger models suffer more, meaning top-ranked models are often just those lucky with the wording. This fragility persists across GSM8K, MMLU, and MATH-Hard, with different models agreeing on which rephrasings cause the most errors.
More from Research
- Paper: Evidence from Nationwide Generative AI Rollout in Pakistan's Courts — soumitrashukla9 · 2026-08-16
- Starfield Fauna dataset released with 20k images — eccLykta · 2026-08-16
- Information Abundance Paradox: Long-context training undermines parametric knowledge — DanielKhashabi · 2026-08-16
- Paper: miRNAs modulate bioelectrical regionalization in multicellular aggregates — drmichaellevin · 2026-08-16
- Intern-S2-Mobius Decouples Knowledge and Reasoning for 4x Speed — jiqizhixin · 2026-08-16
- Position Paper: Teaching as the Grand Challenge for Theory of Mind in AI — atilimgunes · 2026-08-16