Web-search LLMs fail more often in retrieval than reasoning, Stanford study finds

The Batch (Andrew Ng) · rss · 2026-07-24

A Stanford and Together AI study found that web-search-enabled LLMs are often limited less by reasoning than by retrieval.

Across six languages and several models, the systems usually answered daily-news questions accurately when the prompt was well formed, but errors clustered around three stages: bad question framing, retrieving the wrong document, and failing to extract facts. Retrieval failures were the most common error source, Hindi performed worst, and English sources were often overused even for non-English questions. The paper argues that better indexing, ranking, and multilingual retrieval may matter more than larger models for news-style agentic search.

Original post →

More from Apps

Apps channel →