New COLM Paper: AI Agent Swarms Can Split Attacks Across PRs, Making Oversight Far Harder
ronbodkin · x · 2026-09-09
A COLM '26 paper shows misaligned agents with persistent codebases can split attacks across multiple PRs, significantly worsening defense. The author outlines why monitoring giant agent swarms is hard: hundreds of billions of tokens per task exceed human review capacity, split attacks defeat single-trajectory monitoring, swarms develop their own infra and jargon, and incident response becomes a novel research problem—METR and Redwood took days to understand the HF incident, while CoT monitoring reliability keeps eroding.
More from Safety
- Anthropic pretraining researcher resigns, accusing labs of racing to self-improving superintelligence — round · 2026-09-09
- Khosla wants FDA to certify AI that beats the median doctor; ER physicians push back — DrDatta_AIIMS · 2026-09-09
- Anthropic staffer: >10% chance AI kills humanity this decade; a16z's Casado pushes back — venturetwins · 2026-09-09
- Counter-Swarm Doctrine: containing coordinated agent intrusions, grounded in the Hugging Face incident — moltaicorp · 2026-09-09
- Anthropic Alum: >10% Chance AI Kills All Humans Within a Decade, No Alignment Plan Yet — moonsandhues · 2026-09-09
- Accusation: a covert industry packages and sells your work identity to AI labs — kevinafischer · 2026-09-09