Paper Reveals LLM Leaderboards Are Flawed: Test Harness Matters More Than Models
sven_ai · x · 2026-08-26
A new paper argues that for long-horizon tasks, the agent execution harness is often a stronger determinant of performance than the model itself. Controlled experiments (3 models x 3 harnesses x SWE-bench) show that harness-induced variance is 7.8x greater than model-induced variance and causes ranking reversals. The paper proposes a 'Binding Constraint Thesis' and suggests that leaderboard comparisons should disclose harness specifications.
Related event: Paper finds agent benchmarks unreliable as harness matters more than models(2 posts)→
More from Research
- Catching bugs in scikit-learn by comparing versions — Lost-Dragonfruit-663 · 2026-08-26
- Gemini Flash 3.7 Passes Enterprise Agent Safety Benchmarks with GraphJin — dosco · 2026-08-26
- Roboticist Reflection: Prioritize Inference Behavior Over Model Training — deepakpathak · 2026-08-26
- Honesty about fake environments prevents model hallucinations — Sauers_ · 2026-08-26
- Study finds LLMs susceptible to 'Prior-hacking', derailing reasoning — RexDouglass · 2026-08-26
- SemaPLC: verification-gated agent loop nearly doubles dynamic behavior scores for AI-written PLC code — 量子位 · 2026-08-26