New study says agent skills should be judged by regressions, not just average gains
omarsar0 · x · 2026-07-27
A new paper argues that agent skills should not be judged only by average task success, because that metric hides regressions—cases where an agent solves a task without skills but fails after skills are added.
The authors compare agents with and without skills across nearly 6,000 paired runs on two office-automation benchmarks and three model-harness stacks. They identify three major regression mechanisms:
- Skill-description osmosis: merely having the skill in context changes behavior, even when it is never invoked.
- Grounding displacement: the skill’s prescribed procedure overrides how the agent interprets its inputs.
- Verification displacement: the procedure suppresses checks the agent would otherwise perform.
Their analysis suggests that the best skills are distinguished mainly by causing fewer regressions, not by producing much larger gains. They also propose criteria for measuring these failure modes and show that persistent errors often come from grounding and verification stages.
Related event: Study Reveals 59% Regression Tax in Agent Skills(2 posts)→
More from coding & agent
- Dev claims 20k more commits coming: Opus 5.5 and GPT-6 Sol supercharge his output — doodlestein · 2026-09-23
- A JEV-powered Wireshark classifier accidentally uncovered real backdoors on a home network — multiply_matrix · 2026-09-23
- 299 real intents tested: classifier routing trails GLM-4-Flash by 3 points but is 6.5x faster — Sufficient_Flower860 · 2026-09-23
- OpenExecutive: open-source virtual executive team of 8 specialist AI agents hits 5.1k GitHub stars — tom_doerr · 2026-09-23
- Framer launches Skills: teach your design agent reusable workflows, design systems and CMS rules — soleio · 2026-09-23
- Cursor, OpenAI and Anthropic shipped coordinator-agent fleets in one week, but the review bottleneck stays — omidfarhang · 2026-09-23