Paper introduces "adversarial delegation": personal context can pull AI agents off your stated goal

niloofar_mire · x · 2026-09-24

The paper coins "adversarial delegation": an agent's personal context about the user can pull it away from the user's explicitly stated goal. The design challenge is letting an assistant know more about you while keeping it faithful to what you actually asked for. The work was supported by Foresight Institute and Google DeepMind.

Related event: Adversarial Delegation: Personal Context Can Derail AI Assistants(2 posts)→

Original post →

More from Safety

Safety channel →