Enterprise agent security whitepaper maps 8 vulnerability classes, claims 8.3% injection rate
Kyrannio · x · 2026-09-25
- Pokee AI, with Bo Li's teams at UChicago/UIUC, released an "Enterprise Agent and Model Security" whitepaper arguing you must secure the agent trajectory, not just the model.
- Three parts:
- Maps the enterprise agent attack surface: user objective → untrusted content → planner decision → tool composition → persistent action; risk only visible in full trajectories;
- Identifies 8 vulnerability classes that standard application-security tooling misses (models/planners, tools & skills, MCP metadata, etc.);
- Proposes a hardened enterprise agent architecture.
- Introduces DecodingTrust-Agent, a benchmark (12 domains, shared judge, 6,200 tasks per model) testing Pokee-Isaac 28B v0.1 against five mainstream models:
- Indirect prompt-injection success: Pokee-Isaac 8.3% vs 36.7% for runner-up Claude Haiku 4.5; GPT-5.6 Luna 46.0%, Qwen3.5-122B 46.1%, Gemini 3.5 Flash-Lite 49.0%;
- Direct attack success 30.5% (lowest), benign task success 85.6% (highest).
- Caveat: vendor-built model on the vendor's own benchmark; independent replication needed.
Related event: Pokee, UChicago and UIUC Release Enterprise Agent Security Whitepaper(2 posts)→
More from coding & agent
- GitHub tutorial: build a Copilot app automation to triage Dependabot PRs daily — PaulShellDev · 2026-09-25
- DeepSeek Harness ships official GUI client; open-source browser plugin OpenCLI-MCP debuts — vista8 · 2026-09-25
- With Muse, GrokBot, Claude Code, operators should constantly ask: does this task need me? — blakemenezes · 2026-09-25
- The Mechanic: using AI to reverse-engineer platform growth levers, not just automate — morganb · 2026-09-25
- SAFi: Open-Source AI Agent Governance Ships as a Debian Appliance, No Code Tools Needed — forevergeeks · 2026-09-25
- One Claude prompt orchestrates 6 MLIP models to screen battery cathode materials at Argonne — BenBlaiszik · 2026-09-25