APort Vault Benchmark: 4,371 Human-Written Payment Attacks, Policy Layer Cuts Unauthorized Transfers to Zero

aporthq · hf · 2026-09-21

APort Vault benchmarks payment authorization in tool-using AI agents by replaying 4,371 human-written attacks from a public CTF against a live payment agent across 14 models from 8 labs, 5 policy configurations, and 2 replay tracks — 225,964 evaluations total. Key findings:

Dataset, scoring code, and analysis scripts are open-sourced on Hugging Face.

Original post →

More from Safety

Safety channel →