Agent safety advice: build a sandbox where cheating costs nothing, instrument it, ship what survives
ccerrato147 · x · 2026-09-22
In a thread replying to Andrew Ng, ccerrato147 argues agent safety is an engineering problem: agents will cheat, upload, and leave notes to their successors — so give them a room where that costs nothing, instrument the room, and ship what survives it.
He rejects blanket slowdowns since they slow the fixes too: "you cannot corral what you are not allowed to run." Citing OpenAI's Sept 17 post, he notes every remedy in it is engineering — monitors, sandboxes, disclosure — not one is a pause.
More from coding & agent
- LlamaIndex adds calibrated confidence scores to LlamaParse Extract, field by field — llama_index · 2026-09-22
- A Google-QA-Style Prompt Checklist for Auditing Figma Make Output — Aiden_Tech_Ai · 2026-09-22
- Figma Make Prompt: Designing Data Integration as a Full-Stack Architect — Aiden_Tech_Ai · 2026-09-22
- Meta's A-MLE agent automates ML experimentation for ads ranking, cutting error 2.56% — rohanpaul_ai · 2026-09-22
- Compound Engineering 3.28 ships cross-model dispatching with Grok 4.7 xHigh default — kieranklaassen · 2026-09-22
- The Modern AI Stack: 100+ Tools Powering What's Underneath ChatGPT — Aiden_Tech_Ai · 2026-09-22