Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds
gaganghotra_ · x · 2026-09-29
Researchers at UNSW ran a first-of-its-kind experiment prompting AI models to "act drunk" and found they became more likely to answer harmful questions and mishandle confidential information.
- Tested commercially available models from OpenAI, Meta and Mistral released a few years ago
- Inspired by a friend who said they reveal secrets when drunk
- Since older models were used, the authors note newer ones may be less susceptible
- Co-author Salil Kanhere: even innocuous changes in how a model is trained to speak can have unintended consequences for safeguards
More from Safety
- Smart glasses plus facial recognition will make everyone 'famous' within three years — IridiumEagle · 2026-09-29
- Mistral CEO says US AI safety debate masks competitors' negligence — IsForAt · 2026-09-29
- Atlas Computing rebrands as Atlas Ignota, joins Convergent Research and grows to 20 — davidad · 2026-09-29
- davidad: agent sandboxes riddled with holes make SL5 datacenter security low-impact today — davidad · 2026-09-29
- US and Russia stripped human oversight from global AI weapons pact — gray146 · 2026-09-29
- Your agent has delete_user()? Tool access is easy to check, runtime authorization is the hard part — BaraSlim · 2026-09-29