AdaGuard: Adaptive guard models for LLM agents under user-defined policies

Yunhao Feng · hf · 2026-09-29

The paper introduces AdaGuard, a family of 0.6B/4B/8B guard models that assess agent trajectories against user-supplied policies at inference time.

Original post →

More from Safety

Safety channel →