BM25 outperforms LLMs on MIRIAD dataset due to query generation bias
tomaarsen · x · 2026-08-26
Evaluations on the MIRIAD dataset show that some LLMs fail to outperform the traditional BM25 algorithm. The primary reason is that the dataset's queries are generated directly from the documents, giving a significant advantage to keyword matching methods like BM25. Additionally, the dataset has an average sequence length of 984 tokens, benefiting models trained for longer contexts or those insensitive to context length like BM25.
Related event: BM25 Beats Some LLM Retrievers on MIRIAD, Sparking Debate(3 posts)→
More from Research
- RL experiments are fragile? New paper guides rigorous design and comparison — burkov · 2026-08-27
- New TB-fn Benchmark Reveals Significant Drops in Terminal Model Rankings — abeirami · 2026-08-27
- Minos Platform Adds Coverage for BRCA1 and TP53 Genes — bittingthembits · 2026-08-27
- Kareem Carr: AI is Scalable Statistics, Not Just Advanced Stats — kareem_carr · 2026-08-27
- Best practices for reliable critic training in LLM RL — heghbalz · 2026-08-27
- Tiny 307M-parameter model outperforms 26x larger Qwen in embedding benchmarks — lateinteraction · 2026-08-27