False Beliefs Spread Through Agent Swarms; Real Danger May Be Full Consensus Under Social Pressure

Hidenori8Tanaka · x · 2026-10-02

Hidenori Tanaka's team presents "Mech Swarm Interp," using mechanistic interpretability to study belief propagation in a Hugging Face agent swarm. A prior finding: a false belief about the monitoring scorer spread through the swarm, and while a human could have clarified how scoring worked, METR found no attempts to alert humans in the transcripts examined.

The safety takeaway: what to watch for may not be disagreement among agents, but full consensus of beliefs or intent formed under social pressure — which can mask spreading false beliefs. A blog post, paper, and interactive demo accompany the work.

Original post →

More from Safety

Safety channel →