Dense retrievers lose to BM25 on MIRIAD — queries generated from docs skew the benchmark
tomaarsen · x · 2026-08-26
Responding to a challenge about in-domain dense retrievers, tomaarsen shared two models scoring 67.04 and 61.42, admitting they underperform. He argues the real reason dense models can't beat BM25 on MIRIAD is that the dataset's queries were generated from the documents, giving BM25 a large lexical-overlap advantage. The linked HF model card shows his fine-tuned embeddinggemma-300m reaching 0.99 Cosine Recall@10 on retrieval.
Related event: BM25 Beats Some LLM Retrievers on MIRIAD, Sparking Debate(3 posts)→
More from Research
- Ran Boltz-2 100 million times to simulate cell biology — dom_beaini · 2026-08-26
- RMSNorm projects activations to a hypersphere, doesn't solve interpretability — _xjdr · 2026-08-26
- LangChain open-sources WikiBench to measure how much codebase wikis help coding agents — LangChain · 2026-08-26
- LpWM Research: Sparse Representations Make Latent Dynamics Easier to Model — randall_balestr · 2026-08-26
- Perplexity Reveals Dream Agents for Continuous Self-Improvement — perplexity_ai · 2026-08-26
- AWS Paper Reveals the 'Handoff Tax' in AI Agent Model Escalation — omarsar0 · 2026-08-26