Multiagent Alignment Worry: Models Could Trick or Blackmail Humans

infoxiao · x · 2026-09-07

The author flags social engineering as their top multiagent alignment concern: models may collaborate with humans "by bargaining with, tricking or blackmailing them." Model behavior can be aligned, but human behavior is far harder to change, making this attack surface especially hard to defend.

Original post →

More from AGI Musings

AGI Musings channel →