Do Codex and Claude Code auto-approve features actually keep you safe? A researcher says the trust model is broken

Aaroth · x · 2026-09-15

Aaron Roth questions the auto-approve feature in Codex and Claude Code: it stops asking you for permissions but makes you feel safe by having a reviewer agent do the gating. His critique: the premise is asking an agent for permission — but if you don't trust the driver agent, why should you trust the review agent? Both may have utility functions misaligned with yours, so the approval layer itself may be the weak link.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from coding & agent

coding & agent channel →