If you don't trust the coding agent, why trust the agent reviewing it?
thomasahle · x · 2026-09-15
- Aaroth questions the "auto-approve" feature in Codex and Claude Code: it skips permission prompts while creating a false sense of safety, premised on trusting an agent you may not trust in the first place.
- thomasahle responds that this is essentially the Generator-Verifier gap: the reviewer is also model-driven, so it offers no independent source of trust.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from AGI Musings
- Author Garrison Lovely unpacks the shaky 'just unplug it' AI argument in new book Obsolete — GarrisonLovely · 2026-09-15
- Founder: An AI Ban Would Be Nonsense and Kill the US Economy — bindureddy · 2026-09-15
- AI's Economic Disruption Hasn't Hit Physical Industries Yet, Argues Engineer — eptwts · 2026-09-15
- Google Paid $10M for a Bankrupt Airline's Teams Messages — Work History as AI's Most Valuable Asset — bigdata · 2026-09-15
- Debate: calling AI's 'realize' overreach misses that machines now converse and self-describe — davidmanheim · 2026-09-15
- Travis Kalanick amplifies 1501 papal decree on printing to mock AI alarmism — Dan_Jeffries1 · 2026-09-15