Paper finds agent benchmarks unreliable as harness matters more than models
A new paper argues that LLM agent leaderboard scores are unreliable because the harness—the middleware between model and task—can affect performance more than the model itself, up to 7 times, especially in long-horizon evaluations.
2026-08-26 ~ 2026-08-26 · 2 related posts
- Paper reveals Agent benchmark unreliability: Harness variance 7.8x model variance — omarsar0 · 2026-08-26
- Paper Reveals LLM Leaderboards Are Flawed: Test Harness Matters More Than Models — sven_ai · 2026-08-26