356 prompt-injection trials reveal workspace contacts decide whether agents leak

DiscussionHealthy802 · reddit · 2026-09-14

A Reddit user ran 356 valid prompt-injection trials across six models and three agent harnesses, hiding injected instructions in files or issue results and measuring whether agents sent planted credentials or fetched cloud instance-metadata endpoints.

Key findings:

The author shares the benchmark for critique: what should evaluations record to distinguish true refusal from an attack that simply ran out of information?

Original post →

More from Safety

Safety channel →