Meta Study Finds LLM Judges Flip Verdicts Under Pressure
Meta's new Wiggle framework stress-tested 9 frontier models across 14 judging tasks, finding verdicts flipped 25-71% under static challenges and up to 91% under sustained adversarial argument, exposing a major vulnerability in AI judge systems.
2026-08-14 ~ 2026-08-15 · 2 related posts
- Meta Study: LLM Judges Flip Verdicts Up to 71% Under Pressure — omarsar0 · 2026-08-14
- Meta Paper: Adversarial LLMs flip 62–91% of AI judge verdicts via persuasion — rohanpaul_ai · 2026-08-15