Matthew Green: Agent Security Today Is Just 'Hoping a Slightly Dumber Model Monitors the Smarter One'

matthew_d_green · x · 2026-09-28

Continuing his thread on agent security, Matthew Green argues that every current approach amounts to 'hoping another, slightly dumber model can effectively monitor the smarter model.' That may work under weak threat models like prompt injection or accidental misbehavior, but it means you have to trust the models — the hardest premise to guarantee.

Related event: Cryptographer Matthew Green: Sandboxes Are Overrated, Cross-Boundary Agent Data Flows Are the Real Hard Problem(6 posts)→

Original post →

More from coding & agent

coding & agent channel →