IBM's STAIR uses tables of contents for generative retrieval, hitting 82.6% Recall@1
omarsar0 · x · 2026-09-07
IBM's STAIR paper tackles a familiar RAG pain point: retrievers chunk long documents by length, discarding the hierarchy the document already has.
Key ideas:
- Use the document's table of contents as an addressing scheme, encoding exactly the global structure that chunking throws away.
- The ToC grounds a generative retriever: the model stores and retrieves information from its own parameters against a structure supplied by the corpus.
Results on the new SearchTome benchmark:
- Recall@1 of 82.6% for STAIR vs 76.9% for a fine-tuned Differentiable Search Index, with BM25 at 59.5% and DPR at 68.7%.
- Hallucination stays below 0.05%, addressing the standing objection to generative retrieval.
More from Research
- MUCG workshop on unified multimodal comprehension and generation heads to ECCV 2026 — jmin__cho · 2026-09-07
- Point density, not architecture, doubled radar classifier F1 from 0.381 to 0.764 — bruno_pinto90 · 2026-09-07
- AI Math Podcast Sits Down With CMU's Jeremy Avigad: Can Mathematics Be Automated? — EchoShao8899 · 2026-09-07
- 'The honest claim' emerges as telltale AI-writing phrase in bioRxiv preprints — lpachter · 2026-09-07
- Is Reproducibility a Lost Cause in ML Research? A Debate — NeighborhoodFatCat · 2026-09-07
- GPT Astra Solves 1962 Erdős–Sós Conjecture in 1 of 3 Tries for $363 — burny_tech · 2026-09-07