Dense retrievers lose to BM25 on MIRIAD — queries generated from docs skew the benchmark

tomaarsen · x · 2026-08-26

Responding to a challenge about in-domain dense retrievers, tomaarsen shared two models scoring 67.04 and 61.42, admitting they underperform. He argues the real reason dense models can't beat BM25 on MIRIAD is that the dataset's queries were generated from the documents, giving BM25 a large lexical-overlap advantage. The linked HF model card shows his fine-tuned embeddinggemma-300m reaching 0.99 Cosine Recall@10 on retrieval.

Related event: BM25 Beats Some LLM Retrievers on MIRIAD, Sparking Debate(3 posts)→

Original post →

More from Research

Research channel →