Daydreaming attack: 32 black-box calls per skill steal installable agent skills, beats SigLeak 4x
rohanpaul_ai · x · 2026-09-03
The UC Berkeley paper "Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction" (arXiv:2608.26733) details an execution-only attack on hosted agent skills: the victim is never asked to reveal anything. The attacker adaptively crafts tasks to distinguish hidden behaviors, uses attacker-controlled shadow agents to pick a design, and completes each file from stored victim outputs plus local execution checks.
It formalizes three threat levels (Differential, Trace, Output), focusing on Output where only final responses are visible. Across 7 skills and 4 victim models it recovers 86.8% of capability at Output level, nearly 4x better than SigLeak, producing installable skills with a median 32 victim calls even with disclosure defenses enabled. Hiding files and filtering requests is not enough — the work interface itself needs protection.
Related event: Berkeley Paper Reveals Black-Box Attack Can Steal Hidden Agent Skills(2 posts)→
More from coding & agent
- EarlyEval: SJTU cuts agent evaluation costs by predicting outcomes from intermediate behavior — SJTU · 2026-09-03
- ByteDance Seed's HarnessDev tests whether LLMs can build and evolve their own agent harnesses — ByteDance-Seed · 2026-09-03
- Don't take the AI coworker metaphor literally — agent UIs should hide coordination overhead — davidfromkansas · 2026-09-03
- A dedicated desktop terminal for AI agents: buttons are input, but output needs a screen — Cold_Arm3819 · 2026-09-03
- Anthropic staffer: new model feature live on API, coming to Claude Code within a day — trq212 · 2026-09-03
- Simcha Kosman, Owner of ChatGPT's Secure Sandbox, Hosting an AMA — _clickfix_ · 2026-09-03