mini-RAG open-sourced: fully local pipeline with DeepSeek R1, Ollama, Qdrant hybrid search

tahahussein-4623a412 · reddit · 2026-10-01

The author open-sourced mini-RAG, a fully local RAG app running DeepSeek R1 7B via Ollama, with the interesting part being the retrieval pipeline rather than the chat UI.

Full flow: documents → chunking → BGE-M3 embeddings → dense + sparse retrieval → Qdrant → RRF hybrid search → BGE Reranker → relevance gate → context quality filter → context builder → DeepSeek R1 answering with sources.

The author explicitly avoids "retrieve and blindly send to the LLM", insisting on retrieve → rerank → validate relevance → filter → generate. JWT auth and user-level isolation scope retrieval to the authenticated user instead of the whole vector DB. Open source on GitHub; the author asks the community what to change before calling it production-ready, especially around retrieval quality, reranking, context construction, and local inference performance.

Original post →

More from coding & agent

coding & agent channel →