StepGuard: Step-Level Guardrails with Safety-Utility Balancing for Agents

AI45Research · hf · 2026-08-31

StepGuard is a step-level guard model that audits agent actions before execution. It is trained via automatic trajectory generation and balanced reinforcement learning to reduce attacks while minimizing utility loss.

Original post →

More from Safety

Safety channel →