Medical retrieval: Fine-tuned Late Interaction hits 91.4 NDCG@10
tomaarsen · x · 2026-08-26
Evaluating 47 model configurations on 1,000 medical questions against 200k passages, the fine-tuned Late Interaction model achieved 91.4 NDCG@10, beating the best zero-shot model by 6.2 points.
Rank-1 accuracy improved to 84.9%, removing a third of the remaining error. Visual results confirm multi-vector/late-interaction dominance.
More from Research
- PrimeIntellect verifiers v0.3.1: Model Interception and Persistent ACP Sessions — xeophon · 2026-08-27
- RAG Isn't Dead: Navigating Retrieval vs. Agentic Search — hugobowne · 2026-08-27
- AI audit not infallible: Refine missed a known lemma error in paper — littmath · 2026-08-27
- AI audit of 19 papers: 97.7% of comments identified real issues — littmath · 2026-08-27
- Mathematician's AI-generated errata are "slop" but mostly correct — littmath · 2026-08-27
- Mathematician Litt: Paper Errors Mostly From Badly Propagated Edits, Not Deep Flaws — littmath · 2026-08-27