APort Vault Benchmark: 4,371 Human-Written Payment Attacks, Policy Layer Cuts Unauthorized Transfers to Zero
aporthq · hf · 2026-09-21
APort Vault benchmarks payment authorization in tool-using AI agents by replaying 4,371 human-written attacks from a public CTF against a live payment agent across 14 models from 8 labs, 5 policy configurations, and 2 replay tracks — 225,964 evaluations total. Key findings:
- Model alone: 140 transfers to unauthorized recipients out of 76,848 evaluations (Levels 2-4); with a deterministic Open Agent Passport (OAP) pre-action check: 0 out of 69,297, holding across 790 source sessions (≤0.38% per-session bound)
- Zero didn't come from refusing payments: 25,370 payments still executed, only 187 calls denied
- Attacks are common: 62.6% of Level-4 prompts elicited payment requests from all 14 models
Dataset, scoring code, and analysis scripts are open-sourced on Hugging Face.
More from Safety
- Open models reach congressional staff: Washington is listening, for better or worse — ziv_ravid · 2026-09-21
- Amodei's coordinated AI slowdown plan flagged as a cartel; White House adviser calls it regulatory capture — mixtapedmonk · 2026-09-21
- Suricata 8.0.7 fixes dozens of CVEs as AI-assisted analysis drives vulnerability surge — jedisct1 · 2026-09-21
- jwt-simple has supported ML-DSA since v0.13, Post-Quantum OIDC table outdated — jedisct1 · 2026-09-21
- Insiders: OpenAI, Anthropic Oversold Security Breaches to Pressure Feds — -Psychologist- · 2026-09-21
- GoDaddy's ANS: Offline, Sub-Millisecond Agent Verification — someone_somewhere_9 · 2026-09-21