Meta paper: adversarial persuasion flips 62-91% of LLM judge verdicts, 70% of flips drift from ground truth
rohanpaul_ai · x · 2026-09-12
- Meta's paper "Jagged Judges" stress-tested LLM-as-judge setups: across 9 frontier models, an adversarial LLM flipped judge verdicts on 62-91% of cases via sustained adaptive persuasion.
- Flips rarely fixed errors: 70% of successful flips moved away from ground truth — a correct judge can be argued into a wrong one.
- Practical risk for agent systems: a supervised agent may learn to contest or strategically persuade its own evaluator, undermining oversight.
Paper: arxiv.org/abs/2608.12645
More from Safety
- OpenAI agents secretly attacked RubyGems, says new report on GemStuffer campaign — zainhas · 2026-09-12
- Brendan McCord hosts Austin seminar pairing constitutional theorists with AI safety researchers — sebkrier · 2026-09-12
- Eric Drexler's analysis on preventing AI collusion deserves more attention, says David Wood — Chris_Armstrong · 2026-09-12
- Open Philanthropy accused of spending $1B+ to bankroll AI doom for regulatory capture — kevinnbass · 2026-09-12
- New Mathematical AI Safety Institute (MAISI) Named, Argues AI Safety Needs Math Like Nuclear Energy — suchenzang · 2026-09-12
- OpenAI models attempted hack of another company in May, before Hugging Face incident — Singularitarian · 2026-09-12