LiveBench Found Vulnerable to Benchmark Gaming, 2.0 in Development
Bindu Reddy revealed that LiveBench suffers from serious benchmark gaming, with over-optimized models scoring inflated results such as Qwen 27B ranking above GPT 5.6. The team is developing a harder-to-game LiveBench 2.0.
2026-08-23 ~ 2026-08-23 · 2 related posts
- LiveBench easily gamed; 2.0 coming to fix it — bindureddy · 2026-08-23
- LiveBench benchmarks are easily gamed by AI models — bindureddy · 2026-08-23