Uber's Persona Guardrail lifts agent defense accuracy to 95.9%, cuts false approvals to 4.7%
amaarora · x · 2026-10-07
Uber Engineering published Persona Guardrail, a production-deployed runtime defense framework for agentic AI systems, along with the PAGE benchmark.
- Core idea: Existing guardrails mostly target specific attacks like prompt injection and can't ensure agents stay within intended functionality. Persona Guardrail enforces explicit functional boundaries via semantic allowlist/blocklist specifications, with synchronous input and output validation for customer-facing agents.
- PAGE benchmark: Evaluates function-specific guardrails across benign, adversarial, and out-of-domain interactions on both user and agent turns.
- Results: Versus a generic LLM-based guardrail, overall accuracy improves from 85.7% to 95.9%, out-of-domain detection jumps from 57.3% to 93.5%, and the false-approved rate drops from 25.0% to 4.7% — all within its latency budget.
More from coding & agent
- Hermes is getting a local video editor: 42 FFmpeg scripts, no cloud, no API key — Teknium · 2026-10-07
- ykdojo: Code review is dead — what needs reviewing now is the review process itself — ykdojo · 2026-10-07
- Ofir Press: OpenAI's math moment will hit coding in 6-18 months — TimothyDuignan · 2026-10-07
- Princeton researcher: OpenAI's math breakthrough will hit coding in 6-18 months — brianryhuang · 2026-10-07
- Wand raises $7.7M to build pay-per-call "OpenRouter for agent tools" with 2,500 APIs — SimplyAnnisa · 2026-10-07
- GitHub Copilot CLI v1.0.93 rolls out command sandboxing to all users — copilot-cli-release-app[bot] · 2026-10-07