Paper finds agent benchmarks unreliable as harness matters more than models

A new paper argues that LLM agent leaderboard scores are unreliable because the harness—the middleware between model and task—can affect performance more than the model itself, up to 7 times, especially in long-horizon evaluations.

2026-08-26 ~ 2026-08-26 · 2 related posts