mini-RAG open-sourced: fully local pipeline with DeepSeek R1, Ollama, Qdrant hybrid search
tahahussein-4623a412 · reddit · 2026-10-01
The author open-sourced mini-RAG, a fully local RAG app running DeepSeek R1 7B via Ollama, with the interesting part being the retrieval pipeline rather than the chat UI.
Full flow: documents → chunking → BGE-M3 embeddings → dense + sparse retrieval → Qdrant → RRF hybrid search → BGE Reranker → relevance gate → context quality filter → context builder → DeepSeek R1 answering with sources.
The author explicitly avoids "retrieve and blindly send to the LLM", insisting on retrieve → rerank → validate relevance → filter → generate. JWT auth and user-level isolation scope retrieval to the authenticated user instead of the whole vector DB. Open source on GitHub; the author asks the community what to change before calling it production-ready, especially around retrieval quality, reranking, context construction, and local inference performance.
More from coding & agent
- One prompt, 30 minutes: NVIDIA VSS Blueprint 3.3 builds production-line vision AI agents — NVIDIAAI · 2026-10-01
- Developer runs decision models locally on Ollama's new Nimble for sensitive data — ollama · 2026-10-01
- Google AI proposes RRSI to stop recursive self-improving agents from overfitting benchmarks — burkov · 2026-10-01
- DevDay demo: Codex builds Minecraft for 30-year-old Game Boy hardware via ModRetro plugin — pvncher · 2026-10-01
- Full Talking Video From One Image in 30 Minutes for ~$5 — gorkem · 2026-10-01
- tldraw launches ChatGPT plugin: sketch your app's logic and Codex builds it — DavidKPiano · 2026-10-01