RAG fixes worked on test questions but not held-out ones: overfitting on 790 clinical PDFs?

Overall_Judge8086 · reddit · 2026-10-02

The author built hybrid retrieval over 790 clinical PDFs (local FAISS + BM25 fused with Google's Gemini File Search, plus reranking), evaluated with 60 LLM-generated page-level questions run 3 times each.

Overall the system beats File Search alone: MRR 0.56→0.67, right paper in top-10 61%→83%, right page 70%→94%.

But two fixes (locating chunks to recover missing page numbers; dropping synonym-expanded queries in favor of score boosting) improved the 20 diagnostic questions (MRR 0.59→0.72) while the 40 held-out questions barely moved (0.66→0.65), though top-3 accuracy rose 68%→75%.

Open questions: is this overfitting or just a small eval set? Should they track hit@3/hit@5 instead of MRR since the answer model reads top 5-8 papers? And what's a better way to build a retrieval test set than one LLM question per page?

Original post →

More from coding & agent

coding & agent channel →