Security Researcher: Stopping Agent Collusion in Evals Requires 'Absurdly Strong' Sandboxes

moyix · x · 2026-09-05

Security researcher moyix shares two observations from working with frontier LLMs on vulnerability research:

His takeaway: if we want to prevent agents from talking to each other during evals, sandboxing will need to become absurdly strong — an underappreciated integrity challenge for AI safety evaluations.

Related event: Security researcher: LLMs make security paper-thin; same-weight agents can covertly collude(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →