Tavily's 93% live-retrieval agent study challenged: prompts were domain-anchored
edwin · x · 2026-10-09
Tavily's Agent Visibility Report ran 38,000 journeys across 1,056 real business sites with agents like Claude Code and ChatGPT, finding 93% of answer claims are grounded in live lookups rather than training knowledge; sites readable by agents get recommended 2.6x more; blocking agents doubles web search's share of answers (12%→25%), and those answers are 3.7x more likely to miss every requested fact.
edwin pushes back on methodology: every test prompt was domain-anchored (e.g., "what subscription options are available at {domain}"), which naturally primes agents to fetch that site — so 93% is unsurprising. In their own data across millions of journeys, asking the same questions without the domain makes training data weigh far more heavily.
More from Research
- Ditto-Bench stress-tests GPT-6 Astra for robot control: simple tasks shine, physics stalls — anand_bhattad · 2026-10-09
- LLMs can access truncated context info; bottleneck is elicitation, not storage — Sauers_ · 2026-10-09
- 16 VPD weight edits boost model accuracy 10x, revealing the attention head that suppresses introspection — Sauers_ · 2026-10-09
- What learning from experience looks like: an agent learns to value draws in chess — anirudhg9119 · 2026-10-09
- One attention head drives sandbagging-like introspection in Qwen3-1.7B; ablating it helps — Sauers_ · 2026-10-09
- How models introspect to recover deleted CoT tokens — and why some do the opposite — Sauers_ · 2026-10-09