Vector Institute shows hyphenating entities with doc titles fixes RAG's lost-context problem
VectorInst · x · 2026-10-09
Vector Institute's engineering team details a knowledge-graph-based RAG pipeline that tackles lost document context.
- Key trick: hyphenate entities with their source document titles before embedding, preserving document-level hierarchies that standard chunking destroys
- Most RAG systems treat text as disconnected blocks — good at retrieving relevant-sounding paragraphs, bad at connecting facts scattered across pages or quarters
- Uses LangChain's LLMGraphTransformer to extract explicit entities and relationships; full pipeline, evaluation, and open-source code in the technical post
Related event: Knowledge Graph RAG Boosts Multi-Hop Retrieval Accuracy by 54%(3 posts)→
More from Research
- EA-VAE paper in IEEE T-PAMI fixes systematic uncertainty failures in VAEs — enzoferrante · 2026-10-10
- OpenAI reportedly solved 92 of the 500 most important open math problems in one GitHub push — altryne · 2026-10-10
- Quanta asks: is AI the end of math as we know it? — burny_tech · 2026-10-10
- AI math is teleportation to a foggy summit: Cepelewicz's striking mountain metaphor — burny_tech · 2026-10-10
- ICML paper debunks Platonic Representation convergence, proposes Aristotelian view — phillip_isola · 2026-10-10
- Isola lays out the three main pushbacks to the PRH narrative his new paper answers — phillip_isola · 2026-10-10