Study: LLMs Face Attention Dilution and Declining Retrieval in Million-Token Contexts

_reachsumit · x · 2026-07-03

The paper "Can Language Models Actually Retrieve In-Context?" investigates the retrieval capabilities of LLMs within million-token long texts. The study reveals that while LLMs can retrieve information from ultra-long contexts, an "attention dilution" phenomenon occurs as document volume increases, leading to a significant drop in retrieval performance. To address this, the researchers proposed a length-aware fix, offering new optimization strategies for long-context applications.

Original post →

More from Research

Research channel →