CorpusMap: Entity Pages Lift Agentic Search 6.4–11.7 Points and Cut Tokens 34–57%

Follow the Entities: A Corpus Map for Agentic Search

Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam

cs.CL, cs.AI, cs.IR, cs.LG

2026-09-29

CorpusMap turns a corpus into reusable Entity Pages. On 7 models and 3 benchmarks, overall quality beats raw-corpus search by 6.4–11.7 points with 34–57% fewer input tokens.

What problem this solves

Enterprise questions often need evidence split across systems: an approval in email, requirements in an ops doc, latest status in a report. No single file is enough, and the links among files are rarely written down. Those files also sit inside large piles of tickets, drives, and chat logs.

RAG grabs a top-k set and generates. Agentic search lets the model keep searching the whole corpus with tools. The corpus is still a flat file dump. Finding one relevant doc does not tell the agent which other docs talk about the same project. Meeting notes may use an internal codename while the query uses the official name, so searching with the query text misses them. Every new question pays again to rediscover those links. The links themselves are stable: which documents mention the same project or person can be computed offline and reused.

Method

CorpusMap, from KAIST and Microsoft, is an entity-centric navigation layer. It builds a bipartite graph: entities that appear in at least two documents on one side, documents on the other, edges meaning this document mentions this entity. One document can hang off many entities. Documents with no cross-document entity stay reachable by ordinary search.

Each entity is an Entity Page, a dossier: a short overview, key facts tagged with their source, aliases, and paths to every linked document. Source files stay in place. The page is a jump layer, not a replacement.

Offline construction has four stages:

The last three stages also run with GLinker plus templates, no LLM.

At query time the pages are ordinary files, and the agent still uses ls, grep, and cat. BM25 ranks Entity Pages for the question, then ranks documents those pages link to; the agent gets titles and paths, not full text. On EnterpriseRAG-Bench with GPT-5.5, linked-document candidates score 76.60 quality at 88.1k tokens. No candidates still beat Raw Corpus (69.04 vs 65.60) but jump to 496.4k tokens. Extra entity-entity edges move document recall by -2.6 to +0.2 points while making pages 27% to 52% longer.

Results

All three benchmarks use multi-document questions: 80 queries / 2,819 docs on EnterpriseRAG-Bench, 79 / 6,221 on WixQA, 238 / 6,365 on HERB. The main table uses GPT-5.5 and GPT-5.6 Luna, Terra, and Sol, with the same model building the layer and answering. Baselines are raw-corpus search plus Document Page, Group Page, LLM Wiki, and Corpus2Skill. Those four organizers do not reliably beat the raw corpus.

Overall quality averages the per-benchmark quality metrics with equal weight. Relative tokens are the geometric mean of input tokens versus Raw Corpus.

MethodGPT-5.5 quality / tokensLunaTerraSol
Raw Corpus66.11 / 1.00×53.85 / 1.00×65.75 / 1.00×67.69 / 1.00×
LLM Wiki64.49 / 0.63×54.39 / 1.16×58.71 / 1.04×65.15 / 0.48×
Corpus2Skill45.49 / 0.31×37.27 / 0.72×43.86 / 0.66×44.44 / 0.67×
CorpusMap72.55 / 0.43×65.58 / 0.65×72.51 / 0.66×74.68 / 0.43×
Gold Documents79.62 / 0.02×80.00 / 0.04×79.90 / 0.03×79.69 / 0.03×

Versus Raw Corpus, CorpusMap gains 6.45 to 11.74 points (paired bootstrap, Holm-corrected p < 10^{-4}) and uses 34% to 57% fewer tokens. On EnterpriseRAG-Bench with GPT-5.5, document recall goes from 61.62 to 76.17 and correctness from 62.08 to 73.75. The gold-document Oracle still sits about 7 points higher. Corpus2Skill is often the cheapest in tokens and the worst in quality: a topic tree parks documents under a few branches, and the cluster that matters may never be opened.

DeepSeek-V4-Pro: 71.11 vs 68.02. MAI-Thinking-1: 45.07 vs 34.06, but tokens rise from 129.6k to 158.8k. Single-shot retrieval on EnterpriseRAG-Bench with GPT-5.5: BM25 63.66, dense 65.05, HippoRAG 64.66, GraphRAG 47.95, CorpusMap 76.60. A Luna-built map (at most $74.65) lifts GPT-5.5 from 65.60 to 73.59. After amortizing construction, break-even is about 9.0K queries per corpus for GPT-5.5 and 26.3K for Luna. Incremental updates save 34% to 71% of a full rebuild's construction tokens. On the remaining mostly single-document EnterpriseRAG-Bench questions, GPT-5.5 correctness goes from 89.05 to 94.76 and tokens from 250.7k to 171.7k. Qwen3.8-27B on a GLinker map beats Raw Corpus at 2.8k, 10k, and 20k documents, with fewer tokens.

Why it matters

The useful idea is organizing the corpus into reusable jump points, not training another search policy. When the same project and person recur across tickets, mail, wikis, and chat, resolving them once and walking links is closer to something you can ship. A document can sit on several entity paths at once. In the case study, Corpus2Skill parked two specs under one cluster it never opened and reported the document missing; CorpusMap walked a work-item page into the draft, then a Serving Runtime page into v1.

GLinker shows the map can be built without an LLM. A cheap model can build it for a stronger model to read. This is still an incremental gain: the Oracle gap remains, construction is a one-time bill, and low query volume may never pay it back.

Limitations

There is no standalone Limitations section. The ethics statement flags that folding a person or project onto one page can gather privacy that was scattered, and can surface files to people who should not see them. Deployment would need maps split by permission, plus filters.

Scaling stops around 20k distractor documents, well short of the hundreds of thousands across apps in the introduction. In the main table the builder and the answering model are often the same, so the agent reads pages it wrote; the cross-model reuse table only partly separates that. On MAI the map is more accurate and not cheaper, so fewer tokens on average is not true for every model. Token savings also ride on BM25 candidate paths: the graph without candidates is expensive to wander. Answer quality is judged by GPT-5.6 Sol; a DeepSeek re-judge keeps the same ranking on correctness and related metrics, but Kendall's τ on WixQA context recall is only 0.47.

Terms

Source

What people are saying

Related papers

All paper explainers