Berkeley paper: black-box task outputs alone steal 86.8% of hidden agent skill capability

rohanpaul_ai · x · 2026-09-03

A UC Berkeley paper, Daydreaming, shows a harder agent-security problem: legitimate task outputs alone reveal enough to build a portable replacement for a hidden skill — no need to extract the prompt or files.

Key points:

Advice for hosted skill sellers: protect the work interface — trim output detail, expose fewer execution traces, limit and audit adaptive probing.

Related event: Berkeley Paper Reveals Black-Box Attack Can Steal Hidden Agent Skills(2 posts)→

Original post →

More from coding & agent

coding & agent channel →