Traditional nDCG agrees with humans only 53% of the time, RCP-nDCG hits 97%

Nils_Reimers · x · 2026-10-03

Nils Reimers shares human study results questioning traditional retrieval metrics: when an embedding model scores higher on conventional nDCG, humans agree it produces better search results only 53% of the time — a coin flip. With RCP-nDCG, human agreement rises to as much as 97%, suggesting the new metric tracks human judgment far more reliably.

Related event: Cohere's RCP-nDCG Metric Aligns Search Rankings with Human Judgment(2 posts)→

Original post →

More from Models

Models channel →