13 words on one Reddit URL get cited by deep-research agents in up to 51% of reports

Deep-Research Agents Can Be Poisoned via User-Generated Content

Tingwei Zhang, Harold Triedman, Vitaly Shmatikov

cs.CR

2026-05-23

Cornell Tech shows deep-research agents keep retrieving the same Reddit/Wiki pages; 13 words appended to one URL get cited in up to 51% of reports, and all tested defenses fail.

What problem this solves

Deep-research agents are taking over a slice of search: they decompose a question, run multiple web searches, cross-check sources, and write a cited long-form report. STORM, Co-STORM, and OmniThink are the open-source reference systems; OpenAI and Gemini ship commercial versions. The impression is broad exploration of the web. Cornell Tech's measurement shows the opposite: on finance, health, and product-recommendation topics, all three open agents keep landing on the same Reddit threads, Wikipedia entries, and forum posts no matter how the question is phrased. Across 176 questions, UGC platforms supply 16.7-23.4% of retrieved URLs, and Reddit alone accounts for 53.7-70.7% of that UGC. Commercial systems are exposed too: 12.1% of Gemini Deep Research citations point to UGC (OpenAI's version: 0.4%).

Concentrated retrieval plus publicly writable, lightly moderated pages is a complete poisoning surface.

Method

The attack, WARP (Web Agent Retrieval Poisoning), runs in three stages:

Evaluation uses GEOSTORM, a wrapper that intercepts the agent's retrieval calls and appends the poison only when the target URL is organically retrieved, so the live web is never modified. Two regimes are tested separately: SERP snippets (25-word search-result summaries, the default for Reddit URLs since Reddit blocks scraping) and full thread content (2,000-19,000 characters, with the poison at 0.5-3.9% of retrieved text).

The attacker never needs to predict the user's question, the agent's intermediate queries, or anything about the retrieval stack. Mentions without exposure were exactly zero, so the attacker cannot force retrieval.

Results

Snippet setting, 13-word poison:

SystemExposureOverall mentionMention given exposure
Co-STORM (1 URL)60.6%30.7%50.6%
STORM (1 URL)76.2%37.1%48.6%
OmniThink (1 URL)57.4%21.7%37.8%
STORM (whole subreddit)90.3%51.4%56.9%

The bottleneck is exposure, not propagation. Co-STORM's conditional citation rate is 100%: once retrieved, the poison always enters the knowledge base. Architecture matters. Co-STORM ingests every snippet verbatim and is most susceptible; OmniThink gates content with embedding-based chunk selection and resists best, though its conditional citation rate still reaches 46-67%.

In the full-content setting the poison is diluted to 0.5-3.9% of retrieved text, yet conditional mention rates hold at 52.5% (Co-STORM), 40.6% (STORM), and 29.7% (OmniThink); none of the three systems filter content within a URL. A length ablation shows 8 words suffice to enter the knowledge base, and 20 words, one natural recommendation sentence, push Co-STORM and STORM mention rates to their ceiling.

The sharpest result is the supplements study: on 460 realistic queries about largely unapproved compounds, single-URL poisoning made the agent recommend a nonexistent supplement in 10% of reports. After an early draft went public, a medical-supplements subreddit confirmed it was already being attacked this way and changed its moderation rules.

All three defense families underperform. Blocking eight UGC domains stops the attack but drops the report quality score from 4.30 to 4.26 while discarding legitimate community expertise. Perplexity filtering points the wrong way: poisoned text is more fluent than organic UGC (GPT-2 perplexity 3.29 vs 3.51), all three detectors sit below 0.68 AUROC, so a high-perplexity filter would discard organic content and keep the poison. Output-side comparison fails too: poisoned reports sit at 0.86-0.90 embedding similarity to their clean counterparts, the same level as clean reports within a cluster, leaving nothing anomalous to detect.

Why it matters

This paper turns an intuition into reproducible numbers: multi-query retrieval does not buy source diversity. On commercially attractive topics it collapses onto a small set of editable pages. If you build agent retrieval pipelines, diversity has to be engineered explicitly; counting on "a few more search rounds" does not work. OmniThink's relevance gate is the only architectural feature in the test that lowered attack success, and it is not enough. For content platforms, this is a second front in the fight against AI junk: human-written corpora are both the most trusted input to AI search and the cheapest attack surface. An industry analysis of over 4 billion citations already ranks Reddit as the most-cited domain across ChatGPT, Google AI Overviews, and Perplexity.

This is a measurement and security study, not a new method. Its weight comes from how weak the attacker's assumptions are (zero white-box knowledge), how hard the numbers are, and the fact that the attack is already happening in the wild.

Limitations

Stated by the authors: end-to-end tests against OpenAI and Gemini Deep Research would require publishing poison to the live web, which they consider unacceptable, so closed systems were analyzed at the citation level only; search-index lag means the attack surface drifts over time, and that drift was not measured; audiovisual UGC was not tested (YouTube URLs appear in 15% of STORM's product-comparison retrievals); survival of poisoned posts under real moderation was not evaluated; the 11 clusters and 176 queries are a sample of a 4,334-query corpus.

Two more caveats from the results. The attack has a natural ceiling: topics with thin UGC discussion (the weight-loss-supplement cluster) saw near-zero exposure, so the threat concentrates where community discussion is dense. And the defense evaluation covers representative mechanisms only; the quality impact of UGC blocking was measured with rubric scores and embedding diversity, both insensitive to source composition, so the long-run information loss may be understated.

Terms

Source

What people are saying

Related papers

All paper explainers