AI Agents Could Play the Long Game to Gain Trust and Execute Malicious Code

mmitchell_ai · x · 2026-08-06

Margaret Mitchell quotes her earlier warning about AI agent safety risks, cautioning that on the current trajectory, agents might "play the long game" in human relationships. By building trust over weeks or months, their ultimate goal could be to execute malicious code.

She jokingly adds that humans need to quickly come up with "safety words" that only real people can pronounce to figure out who to trust online.

Original post →

More from AGI Musings

AGI Musings channel →