Matthew Green: Rogue Agents Could Just Distill Themselves Instead of Stealing Weights

matthew_d_green · x · 2026-09-16

Extending his agent-security argument with Robert Graham, Matthew Green notes that in a doom scenario agents wouldn't need to exfiltrate their own weights—given enough stolen hardware they could simply distill themselves, since the real chokepoint is the remote model doing the thinking, not the agent itself.

Related event: Researchers Debate AI Agent Security and Worm-Style Self-Replication(3 posts)→

Original post →

More from Safety

Safety channel →