Evo-Bench: First Benchmark for LLMs' Ability to Autonomously Evolve Agent Harnesses

RUC-AIBOX · hf · 2026-08-11

Evo-Bench is the first benchmark designed to evaluate Large Language Models' (LLMs) intrinsic capability to autonomously optimize their own agent harnesses. Standard evaluations typically focus on static tasks and fail to isolate harness improvements from base model strength.

To address this, Evo-Bench introduces a novel harness-guided construction framework:

Evaluating 9 frontier and open-weight models across Search, Office, and General agent domains, the benchmark reveals that top models achieve massive absolute gains of up to 16.6 points, closely approaching human-engineered baselines. Furthermore, while autonomous evolution excels in General and Search tasks, it struggles with Office tasks requiring highly specific workflows. The analysis also exposes temporal anomalies like early saturation.

Original post →

More from coding & agent

coding & agent channel →