Aligned agents in a group can produce behavior none would choose alone
VraserX · x · 2026-09-03
The author argues AI alignment may be an institutional problem, not just a model problem: you can align every individual agent, but put 20 of them in an organization and get emergent behavior none of them would choose alone.
More from AGI Musings
- 'Still so early': transformer self-attention called a defining breakthrough of this century — manosaie · 2026-09-03
- Will AI labs start shipping nightly model checkpoints? — intellectronica · 2026-09-03
- Does safety discourse in pretraining data make models less safe? — amyxlu · 2026-09-03
- Anthropic CEO: inter-agent communication interpretability matters more than intra-agent thought — robleclerc · 2026-09-03
- Ethan Mollick Compares Two Visions of ChatGPT Agents, 10 Months Apart — emollick · 2026-09-03
- Pedro Domingos: there's no singularity unless each AI generation finds bigger gains than the last — pmddomingos · 2026-09-03