IBM Study: Benchmark High Scores Depend on Wording, Strong Models More Fragile

omarsar0 · x · 2026-08-16

New IBM research reveals that model performance on benchmarks is heavily dependent on question phrasing. Using BenchDrift to generate meaning-preserving variations of problems, the study finds that weaker models often benefit from rephrasing, while stronger models suffer more, meaning top-ranked models are often just those lucky with the wording. This fragility persists across GSM8K, MMLU, and MATH-Hard, with different models agreeing on which rephrasings cause the most errors.

Original post →

More from Research

Research channel →