Matthew Green: Agent Security Today Is Just 'Hoping a Slightly Dumber Model Monitors the Smarter One'
matthew_d_green · x · 2026-09-28
Continuing his thread on agent security, Matthew Green argues that every current approach amounts to 'hoping another, slightly dumber model can effectively monitor the smarter model.' That may work under weak threat models like prompt injection or accidental misbehavior, but it means you have to trust the models — the hardest premise to guarantee.
More from coding & agent
- Jev seen as fit for high-level robotics orchestration: fast schema beats LLM lag — ai · 2026-09-28
- Apple's free on-device fm paired with decision model Jev beats big-model routing in tests — jasonkneen · 2026-09-28
- Codex builds whole apps but keeps choking on reusing your live Chrome session — ChrisGPT · 2026-09-28
- DeepSeek ships Codex-style agent harness that connects to multiple model providers — ChrisUniverse · 2026-09-28
- Indie dev builds CoIsland, a notch-resident AI monitor with 37 connectors, entirely with Claude Code — DonTizi · 2026-09-28
- Andrew Carr says Astra gave bit-identical, hash-verified outputs on his weekend project — andrew_n_carr · 2026-09-28