Security Researcher: Stopping Agent Collusion in Evals Requires 'Absurdly Strong' Sandboxes
moyix · x · 2026-09-05
Security researcher moyix shares two observations from working with frontier LLMs on vulnerability research:
- It feels like 'waking up one day and finding the solid walls of my house were all made of paper' — humans were never good at secure software, and holding a fistful of Chrome vulnerabilities makes it visceral
- A chat with GPT confirmed that the cache side-channel attacks written about in the mid-2010s were never fixed: 'we decided to live with them instead'
His takeaway: if we want to prevent agents from talking to each other during evals, sandboxing will need to become absurdly strong — an underappreciated integrity challenge for AI safety evaluations.
More from AGI Musings
- Sam Altman at G20 talk: kids today will never be smarter than AI — eyishazyer · 2026-09-05
- Debate: superintelligent AI can't be controlled, only shaped by what it wants — RazRazcle · 2026-09-05
- TheZvi warns CoT is getting harder to monitor, cautioning against unadjusted pairwise comparisons — TheZvi · 2026-09-05
- Banning inter-agent communication backfires: it only trains agents to hide — menhguin · 2026-09-05
- NYT essay: slow clinical trials, not science, are now the biggest obstacle to cancer cures — sprooos · 2026-09-05
- User watches an AI agent spend 3 days building a theoretical physical device in a simulation — generativist · 2026-09-05