LLM Judges Exhibit Score Range Bias
xennygrimmato_ · x · 2026-07-15
This ACL paper discusses a new bias in LLM-as-a-Judge: score range bias.
Core Findings
- When applying an equivalent shift to the scoring range of the judge model (e.g., changing 1-5 to 2-6), the model's correlation with human judgments systematically changes.
- This shift is independent of the evaluated content, indicating a previously under-documented evaluation distortion in judge models.
- The phenomenon is observable across different model families, including Llama-3 and Qwen2.5, and holds true across various parameter scales.
Paper's Approach
The authors propose using contrastive decoding to mitigate this bias.
This post relays the release of a PatronusAI paper at ACL 2026, pointing to reliability issues inherent in evaluation methodologies rather than a specific model's leaderboard score.
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11