Paper Reveals LLM Leaderboards Are Flawed: Test Harness Matters More Than Models

sven_ai · x · 2026-08-26

A new paper argues that for long-horizon tasks, the agent execution harness is often a stronger determinant of performance than the model itself. Controlled experiments (3 models x 3 harnesses x SWE-bench) show that harness-induced variance is 7.8x greater than model-induced variance and causes ranking reversals. The paper proposes a 'Binding Constraint Thesis' and suggests that leaderboard comparisons should disclose harness specifications.

Related event: Paper finds agent benchmarks unreliable as harness matters more than models(2 posts)→

Original post →

More from Research

Research channel →