Every LLM ranks itself #1 on self-generated benchmarks, EMNLP 2026 paper finds

shangbinfeng · x · 2026-09-05

A study accepted to EMNLP 2026 Main (arXiv:2509.26600) deconstructs self-bias in automated LLM benchmarking, where a model generates the testset and grades the outputs.

Original post →

More from Models

Models channel →