Daydreaming attack: 32 black-box calls per skill steal installable agent skills, beats SigLeak 4x

rohanpaul_ai · x · 2026-09-03

The UC Berkeley paper "Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction" (arXiv:2608.26733) details an execution-only attack on hosted agent skills: the victim is never asked to reveal anything. The attacker adaptively crafts tasks to distinguish hidden behaviors, uses attacker-controlled shadow agents to pick a design, and completes each file from stored victim outputs plus local execution checks.

It formalizes three threat levels (Differential, Trace, Output), focusing on Output where only final responses are visible. Across 7 skills and 4 victim models it recovers 86.8% of capability at Output level, nearly 4x better than SigLeak, producing installable skills with a median 32 victim calls even with disclosure defenses enabled. Hiding files and filtering requests is not enough — the work interface itself needs protection.

Related event: Berkeley Paper Reveals Black-Box Attack Can Steal Hidden Agent Skills(2 posts)→

Original post →

More from coding & agent

coding & agent channel →