Do Codex and Claude Code auto-approve features actually keep you safe? A researcher says the trust model is broken
Aaroth · x · 2026-09-15
Aaron Roth questions the auto-approve feature in Codex and Claude Code: it stops asking you for permissions but makes you feel safe by having a reviewer agent do the gating. His critique: the premise is asking an agent for permission — but if you don't trust the driver agent, why should you trust the review agent? Both may have utility functions misaligned with yours, so the approval layer itself may be the weak link.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from coding & agent
- SWE: pre-nerf Opus 4.6 was the goat, and smarter models won't make SWE easier — LouMM · 2026-09-15
- Elyx debuts with a new AI-friendly design file format for designers and agents — michalmalewicz · 2026-09-15
- Hour-long deep dive with an AI agents expert on workflows, skills, and making money — Rasmic · 2026-09-15
- Skill+CLI vs MCP: a developer asks which integration route wins on tokens and reliability — PrinceSauromates · 2026-09-15
- Open-source skill turns one sentence into a single-file web 3D documentary via Claude Code — xiaohu · 2026-09-15
- Hyper3D launches MCP service, letting Codex auto-build 3D models into explainer sites — xiaohu · 2026-09-15