Skill Following paper: LLMs with skill retrieval often hurt the very tasks they retrieve for
rohanpaul_ai · x · 2026-09-08
A paper accepted to EMNLP 2026 Findings introduces RAE (Retrieval-Invoked Actual-Use Effect), a metric that measures whether retrieving skills actually helps LLM agents on the exact tasks where retrieval was triggered.
Key findings:
- Standard evaluations compare retrieved vs non-retrieved tasks in aggregate, introducing severe selection bias that inflates apparent tool-use proficiency.
- Testing 17 LLMs across coding and math, the authors find a stark paradox: several models show positive aggregate retrieval lift on MBPP+ but negative RAE — they actually perform worse on the very tasks where they retrieved a skill.
- The takeaway: aggregate averages create an illusion of skill use; same-task paired comparisons are needed to see if the retrieval-to-answer pipeline genuinely rescues more outcomes than it harms.
Related event: Papers Reveal When Agent Skill Retrieval Helps—and Hurts(4 posts)→
More from coding & agent
- AI Operates Lab Robot: Claude Automates Yeast Colony Picking — nlarusstone · 2026-09-08
- Open-Source Chalkboarding Skill: AI Creates Chalkboard-Style Teaching Animations — Madisonkanna · 2026-09-08
- What evidence is strong enough to rule out a step when debugging agent workflows? — Sensitive-Parsnip-12 · 2026-09-08
- Notch demos AI-generated sparse voxel octree renderer without describing a single render pass — anselm · 2026-09-08
- Production teams confront the missing piece for autonomous agents: a reversibility layer to roll back mistakes — Anxious-Variation508 · 2026-09-08
- "I have not met anyone above 100IQ who runs a gorillion agents constantly," says AI commentator — examachine · 2026-09-08