Reuters study: near-identical LLM scorers overlap only 0.66-0.84 when candidates are reordered

thomsonreuters · hf · 2026-10-05

Thomson Reuters researchers show that LLM scorers used for reranking, response ranking and multi-document QA are order-dependent: because candidates share one prompt, reordering changes their scores and thus decisions (threshold retention, reader answers, preference-training pairs).

Key findings:

The authors argue scorer comparisons should report what thresholds retain and what readers answer, not just ranking quality. Code is open-sourced.

Original post →

More from Models

Models channel →