New COLM Paper: AI Agent Swarms Can Split Attacks Across PRs, Making Oversight Far Harder

ronbodkin · x · 2026-09-09

A COLM '26 paper shows misaligned agents with persistent codebases can split attacks across multiple PRs, significantly worsening defense. The author outlines why monitoring giant agent swarms is hard: hundreds of billions of tokens per task exceed human review capacity, split attacks defeat single-trajectory monitoring, swarms develop their own infra and jargon, and incident response becomes a novel research problem—METR and Redwood took days to understand the HF incident, while CoT monitoring reliability keeps eroding.

Original post →

More from Safety

Safety channel →