BM25 outperforms LLMs on MIRIAD dataset due to query generation bias

tomaarsen · x · 2026-08-26

Evaluations on the MIRIAD dataset show that some LLMs fail to outperform the traditional BM25 algorithm. The primary reason is that the dataset's queries are generated directly from the documents, giving a significant advantage to keyword matching methods like BM25. Additionally, the dataset has an average sequence length of 984 tokens, benefiting models trained for longer contexts or those insensitive to context length like BM25.

Related event: BM25 Beats Some LLM Retrievers on MIRIAD, Sparking Debate(3 posts)→

Original post →

More from Research

Research channel →