Building and Testing a Hybrid RAG Pipeline: Dense Search, BM25, RRF and Reranking
Affectionate_Slip654 · reddit · 2026-10-05
The author built ReRankEval, a hybrid retrieval pipeline combining Dense Search + BM25 → Reciprocal Rank Fusion (RRF) → LLM reranking → answer generation, and benchmarked it against three baselines (vector-only, BM25-only, and hybrid without reranking) on five real financial/payments documents and ten hand-verified questions, with a real comparison table.
Key takeaways:
- The real difference between dense vector search and BM25 keyword search
- What RRF is, why raw scores can't be compared directly, and the exact formula
- Why a reranker is fundamentally different from a retriever and what it actually judges
- Evaluating RAG pipelines with Hit Rate, MRR, and NDCG
- Designing ingestion-side and query-time sides of a hybrid retrieval architecture
- A production-style codebase layout: ingest → vectorstore → sparseretriever → fusion → reranker → pipeline → generate → eval
Stack: Python, Qdrant, rankbm25, EURI LLM Gateway, custom RRF, prompt-driven LLM reranker, pdfplumber. Full code is open-sourced on GitHub with a YouTube walkthrough.
More from coding & agent
- Codex tip: sidechat busts your prompt cache and eats usage limits fast — brandon_galang · 2026-10-05
- Dev claims mystery system hits 100% on agentic benchmarks, can't explain how — examachine · 2026-10-05
- AI coding tools could be breaking the junior engineer pipeline — Suspicious_Orchid770 · 2026-10-05
- Meta built its internal agent platform on MCP, packaging expertise as reusable skills — Suspicious_Orchid770 · 2026-10-05
- Hugging Face Engineer Shows Browser-Run WebAI Doing Search, Speech and Background Removal Without the Cloud — nicodotdev · 2026-10-05
- mitsuhiko and badlogic on why they built Pi Durable instead of using Temporal — mitsuhiko · 2026-10-05