LLM Judges Flip 10-12% of Close Verdicts When You Swap A/B Order, and Bigger Models Aren't More Robust

Beneficial_Use2116 · reddit · 2026-10-04

A metamorphic-testing experiment probed position bias in LLM-as-judge pipelines: swap the order of two answers and see if the verdict flips. Testing claude-haiku-4.5 and claude-sonnet-5 on 21 pairs with 168 shuffles each, easy factual MCQs and clear-winner controls showed 0% flips—confirming the method detects real bias, not noise. But on genuinely close judgment calls, verdicts flipped 12% (Haiku) and 10% (Sonnet), and the larger model was NOT more robust; two pairs flipped on both models. Implication: LLM-based rankings of close calls are partly driven by position, and scaling the judge doesn't fix it. The author notes limitations (small n, one provider) and released the reproducible script (wobbly), inviting tests on GPT-4o/Llama/Qwen judges.

Original post →

More from Research

Research channel →