Evo-Bench: First Benchmark for LLMs' Ability to Autonomously Evolve Agent Harnesses
RUC-AIBOX · hf · 2026-08-11
Evo-Bench is the first benchmark designed to evaluate Large Language Models' (LLMs) intrinsic capability to autonomously optimize their own agent harnesses. Standard evaluations typically focus on static tasks and fail to isolate harness improvements from base model strength.
To address this, Evo-Bench introduces a novel harness-guided construction framework:
- Task Identification: Uses auxiliary-task evolution to pinpoint tasks genuinely sensitive to framework improvements.
- Stratified Splitting: Employs sensitivity-aware stratified splitting to ensure robust cross-suite generalization.
Evaluating 9 frontier and open-weight models across Search, Office, and General agent domains, the benchmark reveals that top models achieve massive absolute gains of up to 16.6 points, closely approaching human-engineered baselines. Furthermore, while autonomous evolution excels in General and Search tasks, it struggles with Office tasks requiring highly specific workflows. The analysis also exposes temporal anomalies like early saturation.
More from coding & agent
- CMU Launches ExploitBench: Testing AI Agents on Real V8 Exploitation — cyb3rops · 2026-08-11
- Zhipu's ZCode Hits 1M Users: GLM Coding Plan Resets Limits, Boosts Multi-Agent Collaboration — jietang · 2026-08-11
- 4x DGX Spark Cluster Achieves 44.6 tok/s on GLM-5.2 for Real Agent Workloads — EAccelerate_42 · 2026-08-11
- Developer Calls for AI Agent Platforms to Expose Cost Control APIs — arthurcolle · 2026-08-11
- Why Are AI Agents Still Fragile in Real-World Workflows? — No_Progress92 · 2026-08-11
- Agent Memory Distillation boosts small LLM agents by up to 27.2% accuracy via hierarchical teacher memory — kaist-ai · 2026-08-11