Web-search LLMs fail more often in retrieval than reasoning, Stanford study finds
The Batch (Andrew Ng) · rss · 2026-07-24
A Stanford and Together AI study found that web-search-enabled LLMs are often limited less by reasoning than by retrieval.
Across six languages and several models, the systems usually answered daily-news questions accurately when the prompt was well formed, but errors clustered around three stages: bad question framing, retrieving the wrong document, and failing to extract facts. Retrieval failures were the most common error source, Hindi performed worst, and English sources were often overused even for non-English questions. The paper argues that better indexing, ranking, and multilingual retrieval may matter more than larger models for news-style agentic search.
More from Apps
- Reddit users debate how much personal financial data to share with ChatGPT — sanbaeva · 2026-07-27
- AI Tool Reshapes Influencer Marketing: 150 Creators Filtered in 15 Minutes — yangyi · 2026-07-27
- Reddit users share how ChatGPT helps with anxiety and daily life — Jaded-Channel-7169 · 2026-07-27
- Open-source browser tool removes GPT Image 2 texture artifacts with a 0.48M model — Parking_Baby_57 · 2026-07-27
- Sentient OS brings proactive computer-use agents to the Mac with an on-device LLM — TechExpert2910 · 2026-07-27
- ChatGPT Sites turns a truck-service company’s 20-page inventory sheet into a mobile tool — Hacebeanbreakfast · 2026-07-27