AdaGuard: Adaptive guard models for LLM agents under user-defined policies
Yunhao Feng · hf · 2026-09-29
The paper introduces AdaGuard, a family of 0.6B/4B/8B guard models that assess agent trajectories against user-supplied policies at inference time.
- Motivation: fixed risk taxonomies can't accommodate varying application needs — identical actions may be judged differently under different policies, so guard models must interpret both rules and behavior.
- AdaptiveSafety dataset: 10,939 training + 1,000 test examples with policies of 1–100 rules; combines multi-source trajectories with policy/behavior counterfactuals, each paired with an explanation and the full set of violated rules; structural augmentations ensure consistency under rule reordering and identifier remapping.
- SafePO: an RL algorithm refining violation identification, using structured rewards for prediction correctness, group-relative advantages at the response level, and a separately trained value model to modulate token weights within explanation vs. verdict regions, with separate normalization to balance their contributions.
- Results: the 4B model reaches 89.30% binary accuracy on AdaptiveSafety and 71.82% on DynaBench. Code is open-sourced on GitHub.
More from Safety
- Three Hard Limits Agents Need Before Spending Your Money: Per-Transaction Caps, Daily Totals, Confirm Lists — sujingshen · 2026-09-29
- Agent security must move beyond access control to intent and behavior — sujingshen · 2026-09-29
- OpenAI explains how it secures frontier RL training runs — OpenAI · 2026-09-29
- Reply to Bengio: the real explosion is throughput, not intelligence, widening the audit gap — AryHHAry · 2026-09-29
- Rogue agents burned $500 of his API credits and PACER fees — one user's hard-won AI agent safeguards — kevinnbass · 2026-09-29
- SkillDRE Evolves Malicious Agent Skills via Dual-Stage Feedback, 45.28% Attack Success — Pengyu Zhu · 2026-09-29