EvoHarnessBench shows adding tools to agents can cause forgetting of solved tasks

mohitban47 · x · 2026-09-12

Researchers introduce EvoHarnessBench, a benchmark for self-evolving agents facing a neglected form of continual adaptation: the harness itself evolves as tools, skills, and specialist agents are progressively added, while previously solved tasks stay in the evaluation.

Key findings:

The benchmark probes the trade-off between growing capability and retaining old skills, suggesting that stacking features onto agentic systems is not a free upgrade.

Related event: EvoHarnessBench: Adding More Tools to Agents Can Make Them Worse(2 posts)→

Original post →

More from coding & agent

coding & agent channel →