Use a second model to approve tool calls and review an agent’s trajectory

corbtt · x · 2026-07-24

The quoted idea proposes a simple safety pattern for AI agents: place another model instance between tool calls to approve actions, review the trajectory so far, and provide feedback to the main model.

The argument is that the reviewer model should have a different prompt and no incentive to satisfy the user’s goal, so it can act as an independent safety layer. The author frames this as reinventing corporate bureaucracy from first principles, but in a good way.

Original post →

More from coding & agent

coding & agent channel →