LiveBench easily gamed; 2.0 coming to fix it
bindureddy · x · 2026-08-23
Bindu Reddy pointed out that agentic coding on LiveBench is easily "benchmaxxed" (gamed), casting doubt on current rankings like Qwen 27B beating GPT 5.6. The team is working on LiveBench 2.0, which will be harder to exploit.
Related event: LiveBench Found Vulnerable to Benchmark Gaming, 2.0 in Development(2 posts)→
More from Models
- Ox Alpha mystery model scores ~63% on full DeepSWE, on par with GPT-5.6 Sol mid — kimmonismus · 2026-08-23
- MiniMax Music Model Noted for Missing Encoder — kalomaze · 2026-08-23
- Comparison: Grok provides wrong info often, Sol excels at challenging assumptions — jdjohnson · 2026-08-23
- Chinese Flash Models Criticized for Over-Reasoning Latency — oran_ge · 2026-08-23
- Grok 4.6 Tops Agentic Tool Use Benchmark for Banking — XFreeze · 2026-08-23
- Experiment proposed: Local Qwen model on Mac vs $10k cloud security scan — natesiggard · 2026-08-23