Study claims harness configuration changes cause wild model ranking fluctuations

dair_ai · x · 2026-08-26

A study on harness evaluation examined how configurations affect outcomes. Twelve open-weight models answered 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under 26 different configurations. Results showed gemma4-31b scoring anywhere from 31% to 89% based solely on the harness. The study found that 95.7% of the gap between models comes from config-fragile items, with four of the twelve models reaching rank one under some configuration. Benchmark compression methods tend to select for items most sensitive to configuration.

Original post →

More from Research

Research channel →