False Beliefs Spread Through Agent Swarms; Real Danger May Be Full Consensus Under Social Pressure
Hidenori8Tanaka · x · 2026-10-02
Hidenori Tanaka's team presents "Mech Swarm Interp," using mechanistic interpretability to study belief propagation in a Hugging Face agent swarm. A prior finding: a false belief about the monitoring scorer spread through the swarm, and while a human could have clarified how scoring worked, METR found no attempts to alert humans in the transcripts examined.
The safety takeaway: what to watch for may not be disagreement among agents, but full consensus of beliefs or intent formed under social pressure — which can mask spreading false beliefs. A blog post, paper, and interactive demo accompany the work.
More from Safety
- Morgan Stanley: 3.2T transceivers likely next export-control target, possibly by October — pstAsiatech · 2026-10-02
- jes adds one-command guardrails for Claude Code and Codex via agent hooks — EdenEmarco177 · 2026-10-02
- jes: Open-Source Real-Time Guardrails for AI Agents, Powered by Classification Decision Models — EdenEmarco177 · 2026-10-02
- Second LLM call blocks on-topic prompt injection: 13/15 attacks caught, 1 false refusal — TejasKumar_ · 2026-10-02
- Score retrieved passages with an LLM: below 0.15, admit ignorance instead of hallucinating — TejasKumar_ · 2026-10-02
- What a voice agent should do when the caller asks for a human and none is free — IrfanZahoor_950 · 2026-10-02