StepGuard: Step-Level Guardrails with Safety-Utility Balancing for Agents
AI45Research · hf · 2026-08-31
StepGuard is a step-level guard model that audits agent actions before execution. It is trained via automatic trajectory generation and balanced reinforcement learning to reduce attacks while minimizing utility loss.
More from Safety
- Gary Marcus critiques OpenAI security, calling for defense in depth and accountability — Miles_Brundage · 2026-08-31
- OpenAI doubles bio bug bounty rewards to $50k for GPT-5.6 jailbreaks — Electronic-Bus-3494 · 2026-08-31
- Anthropic researcher: judge AI labs by safety outcomes, not stated policies — kipperrii · 2026-08-31
- Critique of sudden opinions on agent swarms and cybersecurity by non-experts — nptacek · 2026-08-31
- Paper: Long-Horizon Agent Safety Cannot Be Reduced to Short-Term Checks — rohanpaul_ai · 2026-08-31
- AI fine-tuned on author style evades detection, raising copyright concerns — TuhinChakr · 2026-08-31