Retrieval expert pushes back: out-of-domain dense models losing to BM25 proves nothing

srchvrs · x · 2026-08-26

Retrieval researcher srchvrs challenged: where is a dense retriever trained on the same dataset? The models used were seemingly out-of-domain, so failing to beat BM25 proves little. Context: he had just published a blog fine-tuning a ColBERT-style multi-vector model for medical retrieval with Sentence Transformers, beating every general-purpose retriever he could find after 14.5 hours on a single RTX 3090. The exchange with tomaarsen's reply highlights a methodology issue: how a benchmark's queries are constructed (from documents or not) can systematically favor BM25.

Related event: Finetuning ColBERT on a single RTX 3090 beats general-purpose retrievers in medical search(13 posts)→

Original post →

More from Research

Research channel →