8,135 Experiments Demystify When Agent Skills Work and Fail

An arXiv paper analyzes Agent Skills through 8,135 cross-benchmark experiments, finding that distilling experience into SKILL.md outperforms workflow memory by 6.06 points, as gains come from stable execution rather than accumulated experience.

2026-09-08 ~ 2026-09-08 · 2 related posts