Salesforce releases EvoHarnessBench to test agents against evolving tool harnesses
Salesforce · hf · 2026-09-10
Salesforce has released EvoHarnessBench on Hugging Face, a benchmark evaluating LLM agents under evolving tools, skills, and agent harnesses. The study reveals persistent gaps in retention, adaptation, and 'harness-induced forgetting' — where changes to the harness itself degrade agent performance — addressing a rarely measured problem in real production agent deployments.
More from coding & agent
- Meta's Muse agent is being oversold: 'can buy' is four steps, and only two are done — KamilKad · 2026-09-10
- opc-skills: open-source agent skills for solo founders on Claude Code and Cursor — tom_doerr · 2026-09-10
- tcut: script terminal sessions in TypeScript, render reproducible MP4/GIF/SVG/HTML — samgoodwin89 · 2026-09-10
- Dev builds AI music theory system: GPT-6 composes a full Bach-style fugue in one prompt — HankYeomans · 2026-09-10
- alchemy releases redesigned CLI with reworked profiles and Cloudflare Telemetry support — samgoodwin89 · 2026-09-10
- Where does your Claude Code session live? HQ proposes Iris MCP for reuse — jacob_posel · 2026-09-10