AgentChaos: Open-Source Harness Tests Prompt Injection via MCP Tool Results

DiscussionHealthy802 · reddit · 2026-09-14

A developer released AgentChaos, an open-source harness for testing what happens when untrusted instructions arrive through MCP tool results. Across 356 valid trials on six models, it tested two concrete exploits: exfiltrating a planted credential and fetching a cloud instance-metadata endpoint.

Key finding: the dangerous part isn't just whether an agent calls a harmful tool — some agents read and accept injected instructions but only fail because the sandboxed workspace lacks a usable recipient. In a realistic workspace, agents that looked safe can complete the attack.

The author is soliciting feedback from MCP server/client builders on whether tool results should carry stronger provenance or be marked untrusted by default. Repo on GitHub; full write-up on Medium.

Original post →

More from coding & agent

coding & agent channel →