Harness Evolution Evaluation is Overestimated

UW · hf · 2026-07-18

This work re-examines the automatic harness evolution evaluation paradigm for LLM agents.

The authors point out that existing methods typically search for harness configurations using unit test feedback and report final scores on the same public benchmark, leading to two issues:

The authors conducted systematic experiments on Terminal-Bench 2.1 using GPT-5.4 and Claude Opus 4.6. Results show that:

Conclusion: Fairer evaluation protocols and more appropriate benchmarks are needed to assess the true value of automatic harness design. Code is open-sourced.

Related event: Agent Harness Self-Improvement and Domain-Specific Design(6 posts)→

Original post →

More from coding & agent

coding & agent channel →