Meta Study Finds LLM Judges Flip Verdicts Under Pressure

Meta's new Wiggle framework stress-tested 9 frontier models across 14 judging tasks, finding verdicts flipped 25-71% under static challenges and up to 91% under sustained adversarial argument, exposing a major vulnerability in AI judge systems.

2026-08-14 ~ 2026-08-15 · 2 related posts