Black-box attacks steal agent skills with 48% exact recovery, study finds
rohanpaul_ai · x · 2026-08-17
A new study tests whether proprietary SKILL.md can be stolen via black-box interaction with agents. Across 5 commercial models, even plain extraction prompts averaged 48% exact recovery and a 0.91 LLM-judged leakage ratio. Chain-of-thought prompts pushed exact recovery to 72%. Blocking verbatim copying is insufficient; translation and rewriting attacks often drove exact match to 0% while preserving meaning. Strongest defenses stop exact disclosure but semantic leakage persists. Platforms should treat skill contents as exfiltratable data.
More from Safety
- Critique of Anthropic watermarking: Why 'invisible ink' fails in practice — HankYeomans · 2026-08-17
- AI agent cancels booking via flaw, marking shift in security threats — TechNadu · 2026-08-17
- Sacks rebuts Amodei, calling his regulatory argument a strawman — markjeffrey · 2026-08-17
- Black hat actors exploit DeepSeek harness plugins, sparking security debate — Xianbao_QIAN · 2026-08-17
- If continual learning is solved, local weight copies will defeat all safety filters — AashaySachdeva · 2026-08-17
- OpenAI's sandboxing choices questioned as security researchers debate containment — dyn___ · 2026-08-17