Eval: New Model Hits 91.4 NDCG@10, Leading by 6.2 Points

tomaarsen · x · 2026-08-26

Evaluated against 47 model configurations (dense, sparse, lexical, and multi-vector) on 1,000 held-out medical questions and 200,000 passages. The new model achieves a 91.4 NDCG@10, leading the best zero-shot model of any architecture by 6.2 points.

Related event: Finetuning ColBERT on a single RTX 3090 beats general-purpose retrievers in medical search(13 posts)→

Original post →

More from Apps

Apps channel →