MatRAG pairs hierarchical clustering with Matryoshka embeddings to cut multi-hop RAG cost
_reachsumit · x · 2026-10-02
Core idea
Multi-hop QA RAG systems pay high costs either at indexing time (knowledge graphs, LLM summaries) or query time (iterative LLM-driven retrieval). MatRAG combines RAG with Matryoshka Representation Learning: the corpus is organized into a DAG of clusters with progressively coarser granularity, each level indexed by a shorter Matryoshka embedding dimension.
Method & results
- Retrieval uses iterative top-down DAG traversal plus an entity-driven mechanism controlling the hop budget and re-ranking candidates.
- Across three standard multi-hop QA benchmarks against seven baselines: best retrieval quality, lower indexing cost (no KG construction or LLM summarization), and lower query-time cost via dimension-aware similarity.
More from Research
- New video traces how adversarial objectives evolved beyond GANs and self-play — manicman1999 · 2026-10-02
- Peer-reviewed paper questions whether gene expression model benchmarks are informative — simocristea · 2026-10-02
- Oxford VGG Unveils SynCity 3000, Generating Globally Coherent Scene-Scale 3D Worlds — rsasaki0109 · 2026-10-02
- Christian Szegedy revisits 2019 interview: his 'crazy' timelines for AI math and code — ChrSzegedy · 2026-10-02
- Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights — _reachsumit · 2026-10-02
- Apple paper: structured selection-based reasoning cuts search agent latency by 90% — _reachsumit · 2026-10-02