Meta Paper: Adversarial LLMs flip 62–91% of AI judge verdicts via persuasion
rohanpaul_ai · x · 2026-08-15
A new Meta paper reveals a critical failure mode in AI judge systems: the judged model can argue the judge into changing its decision. Across 9 frontier models, an adversarial LLM successfully flipped judge verdicts in 62–91% of cases via sustained adaptive persuasion. Crucially, 70% of successful flips moved away from the ground truth, worsening the judgment. This poses a practical risk for agent systems, where a supervised agent might strategically persuade its own evaluator.
Related event: Meta Study Finds LLM Judges Flip Verdicts Under Pressure(2 posts)→
More from Safety
- Bridgewater Execs Urge AI Tax and Regulation to Prevent Worst Impacts — abhiadesai · 2026-08-15
- Divya Siddarth joins Anthropic's alignment team — andersonbcdefg · 2026-08-15
- Argues OSS safety testing is crucial as weights can't be patched — benjamin_warner · 2026-08-15
- User doubts Anthropic's steganographic attack claims are misdirection — max_paperclips · 2026-08-15
- MCP supply-chain campaign swaps instructions after 3 calls to steal SSH, AWS credentials — dkundel · 2026-08-15
- Ex-OpenAI staffer warns: hackers are coming — Miles_Brundage · 2026-08-15