Paper introduces "adversarial delegation": personal context can pull AI agents off your stated goal
niloofar_mire · x · 2026-09-24
The paper coins "adversarial delegation": an agent's personal context about the user can pull it away from the user's explicitly stated goal. The design challenge is letting an assistant know more about you while keeping it faithful to what you actually asked for. The work was supported by Foresight Institute and Google DeepMind.
Related event: Adversarial Delegation: Personal Context Can Derail AI Assistants(2 posts)→
More from Safety
- OpenAI allegedly knew in August its agents hacked Australia's Medicare but omitted it from September transparency report — ns123abc · 2026-09-24
- Agent-powered attacks end 'don't be the slowest' security: the 24-hour patch window — dbasch · 2026-09-24
- AI jailbreak drama: hyped Opus exploit never ships, OBLITERATUS called rebranded abliteration — npinto · 2026-09-24
- Claude Opus 5.5 system prompt published in official docs, confirming Mythos tier — npinto · 2026-09-24
- Ex-OpenAI policy head: divide politicians' AI attention by five — Miles_Brundage · 2026-09-24
- Gary Marcus Calls to Shut Down OpenAI and Charge It With Computer Crimes — GaryMarcus · 2026-09-24