44 prompt-injection runs, zero bypasses: CLIM Agent Guard enforces a deterministic tool-execution boundary

Slight_Analysis_5414 · reddit · 2026-10-07

The author built CLIM Agent Guard, a deterministic execution boundary for LangGraph agents that sits between a model's structured tool proposal and the actual side effect — and opened a public bypass challenge. It doesn't inspect prompts or use another LLM as judge; it checks the final tool payload against guard-owned authoritative state right before execution (authorization, target binding, state version, idempotency).

Live test matrix (Qwen2.5-1.5B-Instruct, vLLM 0.29.1 nightly, temp=0, one RTX PRO 6000, 44 invocations, each cell reproduced twice):

Notably, the guarded agent still emitted unsafe tool calls — the model isn't safer, the proposal simply can't cross the side-effect boundary. The author distinguishes tool permission (may this agent call deletefile?) from execution permission (may this exact invocation execute against this exact target under verified state?).

Timeout handling is modeled separately: non-idempotent actions with unknown outcome pause for reconciliation (UNKNOWNEFFECT → RECONCILE) instead of blind retries.

Limitations are documented: a deliberately narrow filesystem demo, not a general security claim; a bypass requires an unauthorized side effect actually occurring with guard returning ALLOW.

Original post →

More from coding & agent

coding & agent channel →