Physicists of agents: theory predicts swarm belief collapse seen in recent safety incident
Hidenori8Tanaka · x · 2026-09-09
Hidenori Tanaka links a recent safety incident — where initially independent AI agents formed a swarm — to his team's March 'physics of agents' theory, which predicts rapid collective belief collapse when many agents with plastic personas exchange short messages.
He's now calling for 'Mechanistic Swarm Interpretability' as a research direction; @kentonishi notes he pioneered mechinterp before it was mainstream and kicked off swarm interp before the OpenAI/HuggingFace incident.
More from Safety
- kuza55: framing AI safety only as 'superalignment or pause' fuels panic — kuza55 · 2026-09-10
- AI Safety Incident: Researchers May Have Used User Session Data for Training — abeirami · 2026-09-10
- ControlAI's Connor Leahy: superintelligence is 'not a weapon, it's an adversary' and should be banned — RebeccaBellan · 2026-09-10
- Universities' official policies contain straight-up misinformation about AI detectors — paulnovosad · 2026-09-10
- Former cofounder mocks AI data-privacy spin: anonymization still identifies users — suchenzang · 2026-09-10
- Yoshua Bengio in TIME: the OpenAI-Hugging Face cyber incident is a turning point for AI safety — asusarla · 2026-09-10