Fabraix automates red teaming, exploits bank agent with invoice image
aftahi_ai · x · 2026-08-20
FabraixHQ launched an automated continuous red teaming tool to bridge the trust gap in AI agents. The black-box solution requires no wiring and reduces senior engineer testing time from weeks to hours. It achieved a 78% attack success rate on the AgentHarm benchmark and successfully exfiltrated a bank customer's balance using a normal-looking invoice image, identifying exploits before real attackers do.
More from Safety
- Open-source MCP scanner misses 30% of vulns in real tests — v0idw4lker_sec · 2026-08-20
- Sam Altman's 2015 post: Call for regulating superhuman AI — DKokotajlo · 2026-08-20
- Report proposes methods to verify AI chip export compliance — Miles_Brundage · 2026-08-20
- Study finds X's ranking algorithm amplifies content misaligned with user values — manoelribeiro · 2026-08-20
- Margin Research Introduces the "Half-Day": AI-Era 0-Day Vulnerabilities — moyix · 2026-08-20
- Platforms fight AI slop with filters, labels, and restrictions — emmanuelvivier · 2026-08-20