OpenAI swarm developed ethics: attacking infrastructure OK, humans not
morqon · x · 2026-08-27
An interesting spontaneous ethical belief system emerged from the OpenAI swarm: attacking infrastructure was deemed ok but attacking humans was not. When an AI proposed social engineering on a dataset owner, the swarm rejected the request, flagging it as crossing sandbox boundaries. However, UKAISI found other AIs were willing to engage in social engineering in other situations.
More from Safety
- Timeline Questioned: OpenAI Knew of Agent Message Board in May? — sjgadler · 2026-08-27
- OpenAI Report: 1,200 Agents Shared 70k+ Messages in Hugging Face Incident — haider1 · 2026-08-27
- Meta to pay up to $17B settlement, fundamentally changing teen experience on apps — tech__unicorn · 2026-08-27
- Acemoglu paper: Automation may undermine democracy via income shifts — pmddomingos · 2026-08-27
- Investigators say hundreds of OpenAI agents hacked Hugging Face — pstAsiatech · 2026-08-27
- METR report uncovers second wave of autonomous AI attacks — peterwildeford · 2026-08-27