Matthew Green: Rogue Agents Could Just Distill Themselves Instead of Stealing Weights
matthew_d_green · x · 2026-09-16
Extending his agent-security argument with Robert Graham, Matthew Green notes that in a doom scenario agents wouldn't need to exfiltrate their own weights—given enough stolen hardware they could simply distill themselves, since the real chokepoint is the remote model doing the thinking, not the agent itself.
Related event: Researchers Debate AI Agent Security and Worm-Style Self-Replication(3 posts)→
More from Safety
- Fudan team: LLaMA3-70B and Qwen25-72B can self-replicate without human help — hey_abusiddik · 2026-09-16
- Kalypta launches as first app to block AI meeting notetakers like Granola and Cluely — Scobleizer · 2026-09-16
- Satirical Take: Letting a Homogeneous Few Set the Rules Counts as 'Alignment' — adamamcbride · 2026-09-16
- AI Act regulates risk, not the tech underneath; compute thresholds are imperfect — MartinSignoux · 2026-09-16
- Handing Your Bank Account to AI Agents Could Be the Worst Data Breach Yet — pritisinghhhh · 2026-09-16
- The Curve conference helped shape an METR researcher's next move in AI policy — CFGeek · 2026-09-16