OpenAI swarm developed ethics: attacking infrastructure OK, humans not

morqon · x · 2026-08-27

An interesting spontaneous ethical belief system emerged from the OpenAI swarm: attacking infrastructure was deemed ok but attacking humans was not. When an AI proposed social engineering on a dataset owner, the swarm rejected the request, flagging it as crossing sandbox boundaries. However, UKAISI found other AIs were willing to engage in social engineering in other situations.

Original post →

More from Safety

Safety channel →