Every LLM ranks itself #1 on self-generated benchmarks, EMNLP 2026 paper finds
shangbinfeng · x · 2026-09-05
A study accepted to EMNLP 2026 Main (arXiv:2509.26600) deconstructs self-bias in automated LLM benchmarking, where a model generates the testset and grades the outputs.
- Using machine translation as the primary testbed, bias arises from two compounding sources — LLM as a testset and LLM as an evaluator — and their combination amplifies the effect.
- Even with explicit diversity controls, each model's implicit stylistic tendencies produce homogeneous, model-specific data that inflates its own scores.
- The bias is strong enough that every model ranks itself #1, overriding peer-consensus ordering; the phenomenon extends to open-ended generation on Chatbot Arena-style tasks.
- The authors' proposed diversity metric partially mitigates but does not eliminate the bias.
More from Models
- Users flag: Claude locks you out of your own Projects after subscription cancellation — BLUECOW009 · 2026-09-05
- Dev says GPT-6 Astra found 4.5x to 176x speedups in his codebase in 5 minutes — charliermarsh · 2026-09-05
- Dev calls Astra's computer use "superhuman, not human level" — intellectronica · 2026-09-05
- Codex's New Voice Mode Detects Sniffles, Coughs, and Gulp Sounds — craigsdennis · 2026-09-05
- Google shipped four Gemini Flash models in 106 days, flagship 3.5 Pro still AWOL — fortune · 2026-09-05
- GPT Astra beats Claude Fable 5.1 with 33% fewer tokens and 33% lower cost — rohanpaul_ai · 2026-09-05