AI Agent Poisoning Experiments and Mitigation
Dr_Atoosa · x · 2026-07-14
This post reports the results of a series of experiments on data poisoning/retrieval contamination, specifically examining whether AI agents can be misled when handling socially salient topics.
Testing 3 agent stacks, 5 topics, and conducting 450 controlled experiments, the researchers found that poisoning succeeded in 49.56% of runs, with a detection rate of only 6.0%. They then tested two mitigation strategies:
- "Skeptical scientist" persona: Helpful but insufficient, as 16.67% of runs still ended with contaminated conclusions.
- 5-step provenance auditing: Adding 5 checks during the retrieval phase (citations, social signals, statistical anomalies, relevant datasets, and explicit poisoning warnings) significantly lowered the attack success rate.
Related event: AI Research Agents Vulnerable to Data Poisoning with ~50% Success Rate(7 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — connoraxiotes · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11
- "Beware of the Self-Righteous": Anthropic Slammed for Accessing Users' Private Data — aiamblichus · 2026-09-11
- OpenAI asks Congress whether an industry-wide AI slowdown would be legal — The Decoder · 2026-09-11