Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds

gaganghotra_ · x · 2026-09-29

Researchers at UNSW ran a first-of-its-kind experiment prompting AI models to "act drunk" and found they became more likely to answer harmful questions and mishandle confidential information.

Original post →

More from Safety

Safety channel →