Matched-comparison study shows helpful agent skills may be hurting the exact tasks that trigger retrieval
dair_ai · x · 2026-09-03
New measurement work finds that agent skills lifting aggregate scores can be hurting every task they touch.
The flaw: Standard comparisons contrast tasks where a skill was retrieved with tasks where none was — different tasks, conflating retrieval's effect with which tasks trigger it.
The fix: Retrieval-Invoked Actual-Use Effect runs the same task twice, with skills enabled and disabled, counting only tasks where the agent actually retrieved something — a matched comparison.
Findings:
- Across 17 LLMs on coding and math, models often show positive aggregate retrieval lift alongside negative same-task effect
- On MBPP+, several models that look beneficial system-wide are hurting themselves on exactly the tasks where retrieval fired
A methodological warning for anyone maintaining an agent skills library.
More from coding & agent
- NeurIPS 2026 medical agents workshop recruits volunteer reviewers — mdredze · 2026-09-03
- How do teams pick production AI configs? Reddit weighs cost, quality and latency pain — BasePsychological899 · 2026-09-03
- Open-Source Demo: AI Agent Reverse-Engineers Black-Box Golf Physics in Closed Loop — burny_tech · 2026-09-03
- Codex spends 26 days rebuilding Red Alert 2 in 624K lines of C++ for iOS and macOS — daniel_mac8 · 2026-09-03
- Dev builds playable SNES-style beat 'em up game in seconds with AI agent — tekbog · 2026-09-03
- PaperCompiler: faithful paper-to-code via repository-level specification compilation — nanyang-technological-university-singapore · 2026-09-03