Retrieval expert pushes back: out-of-domain dense models losing to BM25 proves nothing
srchvrs · x · 2026-08-26
Retrieval researcher srchvrs challenged: where is a dense retriever trained on the same dataset? The models used were seemingly out-of-domain, so failing to beat BM25 proves little. Context: he had just published a blog fine-tuning a ColBERT-style multi-vector model for medical retrieval with Sentence Transformers, beating every general-purpose retriever he could find after 14.5 hours on a single RTX 3090. The exchange with tomaarsen's reply highlights a methodology issue: how a benchmark's queries are constructed (from documents or not) can systematically favor BM25.
More from Research
- PrimeIntellect verifiers v0.3.1: Model Interception and Persistent ACP Sessions — xeophon · 2026-08-27
- RAG Isn't Dead: Navigating Retrieval vs. Agentic Search — hugobowne · 2026-08-27
- Podcast: What happens when you let an AI run a science lab — JMarty97 · 2026-08-27
- AI audit not infallible: Refine missed a known lemma error in paper — littmath · 2026-08-27
- Mathematician Litt: Paper Errors Mostly From Badly Propagated Edits, Not Deep Flaws — littmath · 2026-08-27
- Terence Tao ran all his published papers through AI error-finding scaffolds — littmath · 2026-08-27