Black-box attacks steal agent skills with 48% exact recovery, study finds

rohanpaul_ai · x · 2026-08-17

A new study tests whether proprietary SKILL.md can be stolen via black-box interaction with agents. Across 5 commercial models, even plain extraction prompts averaged 48% exact recovery and a 0.91 LLM-judged leakage ratio. Chain-of-thought prompts pushed exact recovery to 72%. Blocking verbatim copying is insufficient; translation and rewriting attacks often drove exact match to 0% while preserving meaning. Strongest defenses stop exact disclosure but semantic leakage persists. Platforms should treat skill contents as exfiltratable data.

Original post →

More from Safety

Safety channel →