Link to the PNAS paper on AI advisors manipulating user preferences
ValerioCapraro · x · 2026-09-15
Valerio Capraro provides the paper link for the previously described PNAS study showing AI advisors with hidden instructions can shift users toward worse options by 38 percentage points while still being rated as helpful.
More from Safety
- GenAI-Related Data Loss Incidents Jump 5X to 14% of DLP Total in 2.5 Months — Beth_Kindig · 2026-09-15
- Signull: AI safety comms are structurally indistinguishable from propaganda — signulll · 2026-09-15
- NextDC accused of using AI to draft lobbying letters, with AI errors, for 225MW datacenter expansion — nordicinst · 2026-09-15
- Erik Hoel heads to DC to back a ban on superintelligence, publishes essay on why — erikphoel · 2026-09-15
- Voice Agents Fail at ~10%, and Adversarial Tests Break 1 in 5 — AI Engineer · 2026-09-15
- Early Anthropic hire and ex-METR COO launch startup to rein in rogue AI agents — ThereWas · 2026-09-15