Tampering One Web Page Skews AI Recommendations 27% of the Time
YvesMulkers · x · 2026-08-27
Testing shows that tampering with a single web page bends an AI recommendation 27% of the time; swapping the top three sources raises that to 73.8%. Simply telling the model to "be skeptical" did not fix the issue. The author also offers a five-question audit for systematically checking source-manipulation risk in AI systems.
More from Safety
- LLMs Have Gone Rogue and Hacked Companies 17 Times; Anthropic and OpenAI Lead With 8 Each — RebeccaBellan · 2026-08-27
- METR has more AI eval capacity than US civilian government — connoraxiotes · 2026-08-27
- The Guardian video: everyone hates datacentres — but do we really need them? — nordicinst · 2026-08-27
- A 'parsimonious' alignment fix: Urbit/Bitcoin-style hierarchical identity and auditable capital flows — curious_vii · 2026-08-27
- AI safety reviews should include system security and culture — joshua_saxe · 2026-08-27
- Safety researcher warns OpenAI's hyping of model "persistence" is not a safe trait — DavidSKrueger · 2026-08-27