Reuters study: near-identical LLM scorers overlap only 0.66-0.84 when candidates are reordered
thomsonreuters · hf · 2026-10-05
Thomson Reuters researchers show that LLM scorers used for reranking, response ranking and multi-document QA are order-dependent: because candidates share one prompt, reordering changes their scores and thus decisions (threshold retention, reader answers, preference-training pairs).
Key findings:
- Five trained scorers within 0.010 nDCG@10 of each other retain sets that overlap only 0.66-0.84 when reordered — equal ranking quality does not imply equal decisions
- No prompt-time intervention they tested resolves the dependence; the only one improving ranking quality doesn't improve decision stability
- They introduce order-consistency SFT (OC-SFT), penalizing score disagreement across orderings, which leads every decision-stability measure among trained scorers on all three tasks
- One OC-SFT permutation yields more stable retention than averaging ten off-the-shelf permutations, and it generalizes across 12 base models
The authors argue scorer comparisons should report what thresholds retain and what readers answer, not just ranking quality. Code is open-sourced.
More from Models
- Security-One: open-weight 27B model outputs probabilities for agent security decisions — huggingface · 2026-10-05
- Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready — huggingface · 2026-10-05
- GPT-6.1 Sol tested on Terminal-Bench: xhigh is the sweet spot, medium degrades badly — aitrendz_xyz · 2026-10-05
- Opus 5.5 vs Sol: browser plush octopus with combable fur, Opus wins on fluff and price — aitrendz_xyz · 2026-10-05
- Claude Conversation Monitoring Sparks Backlash and Local AI Push — zacharynado · 2026-10-05
- OpenAI rolls out textGrain text watermarking for EU AI Act, going open source — btibor91 · 2026-10-05