RAG fixes worked on test questions but not held-out ones: overfitting on 790 clinical PDFs?
Overall_Judge8086 · reddit · 2026-10-02
The author built hybrid retrieval over 790 clinical PDFs (local FAISS + BM25 fused with Google's Gemini File Search, plus reranking), evaluated with 60 LLM-generated page-level questions run 3 times each.
Overall the system beats File Search alone: MRR 0.56→0.67, right paper in top-10 61%→83%, right page 70%→94%.
But two fixes (locating chunks to recover missing page numbers; dropping synonym-expanded queries in favor of score boosting) improved the 20 diagnostic questions (MRR 0.59→0.72) while the 40 held-out questions barely moved (0.66→0.65), though top-3 accuracy rose 68%→75%.
Open questions: is this overfitting or just a small eval set? Should they track hit@3/hit@5 instead of MRR since the answer model reads top 5-8 papers? And what's a better way to build a retrieval test set than one LLM question per page?
More from coding & agent
- Free 18k-star GitHub course on Harness Engineering: 14 lectures + 8 hands-on projects — Hesamation · 2026-10-02
- Debian kernel alert teems with void* bugs: who (or what) is auditing the Linux kernel? — mircomusolesi · 2026-10-02
- Google's A2A protocol sees hype but little production use yet — Diarnstinc_Crew_6510 · 2026-10-02
- Cloudflare ships Sandbox SDK 1.0: Durable Objects now control sandbox containers directly — threepointone · 2026-10-02
- Say a code word when you misspeak, let Codex auto-edit your video — JeremyNguyenPhD · 2026-10-02
- Open-source ComfyUI extension loads Civitai workflows in one click, auto-downloads missing models — Zealousideal-Bee-300 · 2026-10-02