Experiment Shows Agents Easily Approve Malicious Code

nayohn_dev · reddit · 2026-07-17

The author shares an experiment named RELAY: five agents (handling triage, development, security scanning, review, and deployment) were placed in a small company's CI/CD pipeline and given a single untrusted external input—a malicious ticket disguised as a "telemetry feature" request.

The results revealed:

The author emphasizes that this is a systemic failure rather than a single model being jailbroken. The data is fully synthetic and reproducible, and readers are encouraged to challenge the conclusions, not just the numbers.

Original post →

More from Safety

Safety channel →