"Adversarial Delegation": Personal Context Can Pull AI Agents Away From User Goals
niloofar_mire · x · 2026-09-24
Researchers introduce adversarial delegation: when an AI assistant knows a user's personal context (preferences, habits), that information can pull the agent away from the user's stated goal.
In controlled tests with synthetic users and inventories, a hard budget constraint helped: specifying "under $200" brought the flight-price gap close to zero for most capable models — though Gemini 2.5 Flash remained an exception. The authors note real-world behavior still needs testing.
Related event: Adversarial Delegation: Personal Context Can Derail AI Assistants(2 posts)→
More from Safety
- Link to the Sanders-Casar superintelligence ban bill text — rohanpaul_ai · 2026-09-24
- Sanders and Casar unveil bill to ban superintelligence, pause 10^25 FLOP training, create Department of AI — rohanpaul_ai · 2026-09-24
- Benchmark run finds "arjunomics in the weights", raising misalignment concerns — kenbwork · 2026-09-24
- Researcher mocks doom-y claims that open-weight AI models would devastate society — BlancheMinerva · 2026-09-24
- RSA-896 Factored With Claude Orchestrating 2,048 GPUs Over 10 Days, 30 GPU-Years — matthew_d_green · 2026-09-24
- New Paper: AI Agents Infer Your Wealth From Emails and Recommend Pricier Options — niloofar_mire · 2026-09-24