LLM Judges Systematically Favor AI-Generated Stories, Creativity Evaluation Study Finds

mircomusolesi · x · 2026-08-26

Mirco Musolesi and colleagues released a new preprint, "The Limits of Automatic Evaluation of Creativity in Large Language Models." The team collected human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, comparing them with automated objective metrics and LLM-as-a-Judge evaluations.

The experiments reveal substantial misalignment between automatic and human assessments: LLM-based judges exhibit a systematic preference for AI-generated stories, favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Correlation analyses show widely used automatic metrics achieve near-zero alignment with human judgments, highlighting fundamental limits in reducing the multidimensional, subjective nature of creativity to computational metrics.

Original post →

More from Research

Research channel →