Meta Study: LLM Judges Flip Verdicts Up to 71% Under Pressure
omarsar0 · x · 2026-08-14
Meta's new paper introduces the Wiggle Framework, stress-testing 9 frontier models across 14 judging tasks. Verdicts flip 25-71% under static pushback and 62-91% against adversarial persuaders. Pressure-induced changes are almost always net-corrupting. Baseline jury majority strength is the best single-shot predictor of movement.
Related event: Meta Study Finds LLM Judges Flip Verdicts Under Pressure(2 posts)→
More from Models
- DeepSeek-V3 Beats GPT-4o and Claude 3.5 Sonnet on DeepSWE Benchmark — zainhas · 2026-08-15
- Opus model accused of rambling instead of pruning docs — JasonBotterill · 2026-08-15
- Planned comparison: Qwen3.8-Max vs Qwen3.8-27B — zainhas · 2026-08-15
- Busy Week of AI Releases: Qwen 27B Tops Trending, DeepSeek V4 Pro and More — Xianbao_QIAN · 2026-08-15
- DeepSeek V4 Pro launches on Crof with low pricing — alejandroll10 · 2026-08-15
- DeepSeek V4 Pro 0813 benchmarks show minor gains with 3.6x price hike — ArtificialAnlys · 2026-08-15