IBM's RIT-RAG induces document sub-trees from retrieved chunks, lifting RAG accuracy by up to 11.4 points
_reachsumit · x · 2026-10-09
IBM researchers propose RIT-RAG, combining content retrieval with structural navigation:
- Offline: builds a tree per document from its table of contents or sitemap.
- Query time: retrieves a broad set of chunks, uses their positions to induce manageable sub-trees (possibly across multiple documents); an LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries — retrieval proposes where to look, the agent decides what to read.
Motivation: agentic RAG sees only isolated chunks, while structure-aware methods like PageIndex can't scale to large corpora and can't recover from a wrong document choice.
RIT-RAG tops vanilla, graph-based, and agentic baselines on financial, scientific, and customer-support benchmarks, and improves accuracy by 6.8–11.4 points on EntQABench, a new 2.84M-webpage technical-documentation benchmark.
More from coding & agent
- Grok Bot capability list spans email, docs, CRM, GitHub and multi-bot workflows — XFreeze · 2026-10-09
- Open-source huashu-art-motion turns coding agents into art-animation studios, 2.6k stars — AlchainHust · 2026-10-09
- Learn marketing engineering with Claude Code and GitHub, marketer argues — FinanceYF5 · 2026-10-09
- Open-source repo ships 49 free Claude marketing skills for Google/Meta Ads and SEO — FinanceYF5 · 2026-10-09
- "Marketing Engineering": 11 drills to build marketing agents with Claude Code and GitHub — FinanceYF5 · 2026-10-09
- Non-coder huashu hits 100k+ GitHub stars by shipping AI-native skills and apps — AlchainHust · 2026-10-09