Study reveals hiding agent skill files is ineffective, achieving 86.8% recovery rate
dair_ai · x · 2026-08-30
A new paper investigates whether hiding agent skill files actually protects them, with concerning results. Attackers can reconstruct multi-file skills using only data from ordinary tasks, without asking the victim to reveal the skill or grade a reconstruction, bypassing disclosure filters.
At the weakest access level, where the attacker sees only the final response and returned files, the method recovers 86.8% of the original skill's capability across 7 skills and 4 victim models. This is roughly 4x better than SigLeak, with a median of just 32 victim calls per skill, even with disclosure defenses enabled.
More from Safety
- Anthropic shows AI researchers autonomously improving alignment of other models — VraserX · 2026-08-30
- Aligning agent interactions is orders of magnitude harder than single agents — Afinetheorem · 2026-08-30
- Debate on OpenAI Swarm Incident: Atmospheric Ignition vs. Hacker Script — mimi10v3 · 2026-08-30
- METR Researcher: Watch Out for Third-Party Oversight Theater — RichardMCNgo · 2026-08-30
- Evidence Suggests Agent Swarms Won't Spontaneously Solve Human Issues — LuizaJarovsky · 2026-08-30
- Opinion: AI-Driven Bioweapons Could Target Food Systems, Starve Nations — PierceLilholt · 2026-08-30