LLM Judges Flip 10-12% of Close Verdicts When You Swap A/B Order, and Bigger Models Aren't More Robust
Beneficial_Use2116 · reddit · 2026-10-04
A metamorphic-testing experiment probed position bias in LLM-as-judge pipelines: swap the order of two answers and see if the verdict flips. Testing claude-haiku-4.5 and claude-sonnet-5 on 21 pairs with 168 shuffles each, easy factual MCQs and clear-winner controls showed 0% flips—confirming the method detects real bias, not noise. But on genuinely close judgment calls, verdicts flipped 12% (Haiku) and 10% (Sonnet), and the larger model was NOT more robust; two pairs flipped on both models. Implication: LLM-based rankings of close calls are partly driven by position, and scaling the judge doesn't fix it. The author notes limitations (small n, one provider) and released the reproducible script (wobbly), inviting tests on GPT-4o/Llama/Qwen judges.
More from Research
- New paper finds interpretability tools rarely help agents judge whether LLM behavior explanations are true — Sauers_ · 2026-10-04
- 13 AI Models Play Doctor: All 195 Consults Diagnosed Right, but Safety Set Them Apart — radeon2000 · 2026-10-04
- MuscleMimic: open-source benchmark controls all 354 human muscles, zero-shot on chained movements — TinfoilTricorn · 2026-10-04
- Yacine Calls Out Papers That Beat 'SOTA' by Comparing Against Untuned Baselines — yacineMTB · 2026-10-04
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Microsoft's ActiveSaddler adapts agent harness training scenarios, boosting Pass@1 by up to 7.5 points — dair_ai · 2026-10-04