ToolHazard: A Scalable Framework for Adversarial Security Evaluation of LLM Agents

PekingUniversity · hf · 2026-08-13

ToolHazard is a scalable framework designed to synthesize adversarial environments to test LLM-based agents against indirect prompt injections. It effectively reveals agent vulnerabilities and helps improve defensive alignment strategies.

Original post →

More from Safety

Safety channel →