LLM Judges Systematically Favor AI-Generated Stories, Creativity Evaluation Study Finds
mircomusolesi · x · 2026-08-26
Mirco Musolesi and colleagues released a new preprint, "The Limits of Automatic Evaluation of Creativity in Large Language Models." The team collected human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, comparing them with automated objective metrics and LLM-as-a-Judge evaluations.
The experiments reveal substantial misalignment between automatic and human assessments: LLM-based judges exhibit a systematic preference for AI-generated stories, favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Correlation analyses show widely used automatic metrics achieve near-zero alignment with human judgments, highlighting fundamental limits in reducing the multidimensional, subjective nature of creativity to computational metrics.
More from Research
- Is Work and Play the Same? Weekly Roundup on WebGPU and Strudel — generatecoll · 2026-08-26
- AAAI talk "Where Does Agency Live?" now on YouTube — AnnaCiaunica · 2026-08-26
- Alibaba's CommerceAgentBench: 107 Real E-Commerce Tasks, Top Model Fails 40% — iamfakhrealam · 2026-08-26
- Alibaba Proposes DREAM: Agentic Meta-Control for Industrial Recommenders — alibabagroup · 2026-08-26
- Study Shows Diminishing Returns of Prompt Engineering in Newer LLMs — kalyan_kpl · 2026-08-26
- Long Task Success Rate Doubled: Alibaba's MEA Loop Fixes Context Decay — 大模型之路 · 2026-08-26