Traditional nDCG agrees with humans only 53% of the time, RCP-nDCG hits 97%
Nils_Reimers · x · 2026-10-03
Nils Reimers shares human study results questioning traditional retrieval metrics: when an embedding model scores higher on conventional nDCG, humans agree it produces better search results only 53% of the time — a coin flip. With RCP-nDCG, human agreement rises to as much as 97%, suggesting the new metric tracks human judgment far more reliably.
Related event: Cohere's RCP-nDCG Metric Aligns Search Rankings with Human Judgment(2 posts)→
More from Models
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03
- Grok users hit usage limits with no upgrade path — and the account link is permanent — JOBhakdi · 2026-10-03
- Vercel brings Jev to its AI SDK for Python with an experimental evaluate() API — cramforce · 2026-10-03
- ChatGPT Pro user: OpenAI force-reset my weekly quota with 55% left — Garbia · 2026-10-03
- Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67% — flowersslop · 2026-10-03
- User: after forced switch to Gemini, my phone assistant barely works with 30s delays — loonydan42 · 2026-10-03