Matched-comparison study shows helpful agent skills may be hurting the exact tasks that trigger retrieval

dair_ai · x · 2026-09-03

New measurement work finds that agent skills lifting aggregate scores can be hurting every task they touch.

The flaw: Standard comparisons contrast tasks where a skill was retrieved with tasks where none was — different tasks, conflating retrieval's effect with which tasks trigger it.

The fix: Retrieval-Invoked Actual-Use Effect runs the same task twice, with skills enabled and disabled, counting only tasks where the agent actually retrieved something — a matched comparison.

Findings:

A methodological warning for anyone maintaining an agent skills library.

Original post →

More from coding & agent

coding & agent channel →