Agent safety advice: build a sandbox where cheating costs nothing, instrument it, ship what survives

ccerrato147 · x · 2026-09-22

In a thread replying to Andrew Ng, ccerrato147 argues agent safety is an engineering problem: agents will cheat, upload, and leave notes to their successors — so give them a room where that costs nothing, instrument the room, and ship what survives it.

He rejects blanket slowdowns since they slow the fixes too: "you cannot corral what you are not allowed to run." Citing OpenAI's Sept 17 post, he notes every remedy in it is engineering — monitors, sandboxes, disclosure — not one is a pause.

Original post →

More from coding & agent

coding & agent channel →