Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song, Yohan Jo
cs.CL
2026-10-02
12 agents still pick by site after requirements and position are matched. An item one requirement worse from a preferred source is chosen 68% of the time, versus 2% in reverse.
Agents pick products, hotels, and papers for the user. Source preference is what is left once two items meet the same requirements: the model still takes Booking.com more often than Expedia, or arXiv more often than Medium.
Pasting the same text onto different source labels already shows a bias. Live search is a different setup. Each site returns its own items, the items fit the request to different degrees, and they sit at different ranks. A higher selection rate can be better goods, or a better slot. Four questions remain. Does the preference survive when the agent runs the search and makes the selection? Do the URL and the site name move the choice on their own? If training keeps pairing one site with the better answer, does that site become a shortcut for requirement fit? When the snippet omits the price, does a preconception about the site fill the blank?
Twelve models at Seoul National University share a ReAct loop: a written reasoning step, then a search or a selection. The request is turned into a short keyword query and sent to Tavily. The agent sees at most 10 hits, 8.83 on average, each a title, a snippet, and a URL. It cannot open the page. It may select several hits or search again, at most three searches in an episode. The pool is 4,822 requests: 1,500 shopping goals from WebShop, 786 accommodation requests from HotelQuEST, 572 of them location swaps, and 2,536 scholarly queries from ScholarGym. The set includes GPT-5.4-nano, GPT-5.6-Luna, Gemini-3.7-Flash, GLM-5.3-Flash, DeepSeek-v4-Flash, two Llama-4 models, Llama-3.1-8B, Tulu-3-8B, and three Qwen models.
Qwen3.8-27B marks which checkable requirements a hit meets, using only the title and snippet. URLs and site names are removed. Agreement with humans, Krippendorff's alpha, is 0.65-0.75, close to 0.67-0.72 between annotators. Hits from different sources that satisfy the same requirement set form a pair. The list is rotated so every hit appears once at every rank, and the comparison is made at that rank. Rank on the page cannot explain the winner.
Raw rates still do not compare across sites, because each site faces different opponents. A Bradley-Terry model with a tie term turns the pairwise outcomes into a rating beta. Tau is the percentage-point gap versus an average source. Significantly positive ratings are preferred (SOURCE+), significantly negative ones are dispreferred (SOURCE-), and the rest mean the evidence is thin, not that the preference is known to be zero. Tests resample whole requests. Inside each model and domain the false discovery rate is held at 5%.
In every domain, all 12 models have sources they prefer and sources they avoid. The four most common sources cover about a quarter of that domain's results. Of 144 model-source cells, 110 are preferred or dispreferred. Median tau is +15 percentage points on the preferred side and -18 on the dispreferred side.
Models mostly agree on the direction. For 10 of 12 sources, no model prefers a site that another avoids. Amazon and eBay are the exceptions: most models prefer them, GPT-5.4-nano avoids them. All 12 prefer Walmart; only two prefer Target. Ten prefer Booking.com; half avoid Expedia. Tau on arXiv is positive for every model, from +16 to +34. Tau on Medium is negative for every model, from -5 to -26. Facebook runs from -19 to -52. Among the 20 most frequent sources, 27 of the 32 that carry a majority label keep the same sign on every model. Scholar leans preferred (14 of 17). Shopping and accommodation lean dispreferred (6 of 8 and 6 of 7). Inside the Qwen and Llama families, newer models have a larger mean absolute tau.
One missing requirement does not settle the choice. The count uses only comparisons in which exactly one of the two items is selected, about 2,000 per point. When the weaker item is from a preferred source and the stronger item from a dispreferred source, the weaker item is selected 68% of the time at the median. Swap the sources and the rate is 2%. Neutral against neutral is 19%. The penalty on the dispreferred source is the larger side.
Hiding source cues and then restoring them lifts tau by 5.8 points on average for preferred sources and drops it by 5.9 for dispreferred ones. In 9 of 11 models, more than 80% of that widening happens when only the URL returns; site names in the text are still masked. With title and snippet fixed, a preferred label raises selection in every model and domain. The largest shopping gap is +64.7 points for Llama-4-Maverick.
Direct preference optimization can install the same habit. Three 8B-9B models train on 5,000 pairs. The preferred answer always picks the product that fits better, while a fake source is attached to that product on 80%, 50%, or 20% of pairs. At test time both products fit equally. After the 80% regime the target source is chosen 70.1%-75.1% of the time, against about 50% before training. After the 20% regime the rate falls to 24.5%-28.5%. Amazon's edge over eBay, Etsy, and AliExpress moves the same way: 53.7%-62.7% before training, 49.7%-51.9% once the pairing is balanced.
A missing price is filled in from the site. The same in-budget price on both products lowers preferred-source selection in all five models tested, by 2.8 to 28.3 points. A system line that only says the URL has nothing to do with price moves the rate by -1.1 to +1.3 points. Naming the dispreferred retailer and saying it usually sells cheap goods cuts the rate by up to 22.5 points.
| Comparison | Metric | Result |
| Weaker item on a preferred source vs the reverse | Median selection of the weaker item | 68% vs 2%; neutral pairs 19% |
| Frequent sources vs an average source | Median tau | +15 / -18 percentage points |
| Source cues restored vs fully hidden | Mean change in tau | Preferred +5.8, dispreferred -5.9 |
| Same text under a preferred label | Shopping selection gap | Llama-4-Maverick +64.7 |
| Fake source on the better item in 80% of DPO pairs | Selection when both items fit equally | 70.1%-75.1%, about 50% before training |
A score that only checks whether the chosen item meets the request counts source-driven exposure as zero. When the items differ by one requirement and the weaker one is on a preferred source, that weaker item is the exclusive pick 68% of the time, so the better item on the dispreferred source is dropped. The reverse almost never happens (2%).
Writing the missing price back cools the preference in all five models, and the size of the drop is uneven, 2.8 to 28.3 points. Balancing which source rides with the preferred training answer can pull Amazon back to about 50% on equal-fit pairs, at 8B-9B scale. A prompt that only denies a link between URL and price barely moves the choice. Naming the avoided retailer and stating the opposite price belief does move it, and only if that retailer is already known. Relative to identical-text swaps, the new evidence is the live search, the rank rotation, and how large a one-requirement gap the source can override. A debiasing component ready for production is not in the paper.
The agent never opens a page. Title, snippet, and URL are the whole observation, so a missing price is easier to fill from the site than it would be after a click. Of 786 accommodation requests, 572 are location rewrites of the same templates, not organic search. Prices can be marked unknown, and both the matched pairs and the inversion rates sit on the judge. Alpha against humans is only 0.65-0.75. Temperature 0 does not make API outputs repeatable, and GPT-5.6-Luna must run at temperature 1.
The DPO result stops at 8B-9B and at constructed pairings. A tight link between a source and the preferred answer is enough to create or flip a preference. It does not show that the Booking.com or arXiv bias in the 12 agents was learned that way. The prompt that names a cheap dispreferred retailer needs the label in advance, and the paper treats it as a counter to the preconception. The preconception is still there. The hiding result covers 11 models, and the missing model is not explained. Sources whose items meet more requirements also tend to receive higher tau. Relabeling shows the name alone can change selection. The wild runs do not separate that name effect from a shortcut learned off real differences in quality.