Use a second model to approve tool calls and review an agent’s trajectory
corbtt · x · 2026-07-24
The quoted idea proposes a simple safety pattern for AI agents: place another model instance between tool calls to approve actions, review the trajectory so far, and provide feedback to the main model.
The argument is that the reviewer model should have a different prompt and no incentive to satisfy the user’s goal, so it can act as an independent safety layer. The author frames this as reinventing corporate bureaucracy from first principles, but in a good way.
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11