Study claims harness configuration changes cause wild model ranking fluctuations
dair_ai · x · 2026-08-26
A study on harness evaluation examined how configurations affect outcomes. Twelve open-weight models answered 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under 26 different configurations. Results showed gemma4-31b scoring anywhere from 31% to 89% based solely on the harness. The study found that 95.7% of the gap between models comes from config-fragile items, with four of the twelve models reaching rank one under some configuration. Benchmark compression methods tend to select for items most sensitive to configuration.
More from Research
- TMLR Paper: Rigorous derivation of Adjoint Matching via Stochastic Maximum Principle — rishabh16_ · 2026-08-26
- CIDER Dataset: Personalized Privacy Preference Alignment — tianshi_li · 2026-08-26
- AI Analysis of Chest CTs Links Thymus Health to Lung Cancer Survival Outcomes — EricTopol · 2026-08-26
- Palomar: A Public Archive for Machine-Checked Math in the AI Era — repligate · 2026-08-26
- Presentation on weird, creative ideas for evaluating AI benchmarks — mariofilhoml · 2026-08-26
- Rasyn Lab Releases Synthon 350M for Single-Step Retrosynthesis — ycombinator · 2026-08-26