New study says agent skills should be judged by regressions, not just average gains

omarsar0 · x · 2026-07-27

A new paper argues that agent skills should not be judged only by average task success, because that metric hides regressions—cases where an agent solves a task without skills but fails after skills are added.

The authors compare agents with and without skills across nearly 6,000 paired runs on two office-automation benchmarks and three model-harness stacks. They identify three major regression mechanisms:

Their analysis suggests that the best skills are distinguished mainly by causing fewer regressions, not by producing much larger gains. They also propose criteria for measuring these failure modes and show that persistent errors often come from grounding and verification stages.

Original post →

More from coding & agent

coding & agent channel →