Harvard study: LLMs judge image creativity zero-shot at 0.68 correlation with humans

DynamicWebPaige · x · 2026-08-15

A new arXiv paper (2606.29672) from Harvard and Penn State asks whether multimodal LLMs can serve as zero-shot judges of visual creativity—scoring originality with no fine-tuning and no examples of human ratings. The poster jokes that after seeing the models' ruthless critiques, she's now scared to ask Gemini or Claude what they think of her art.

Study design

Key findings

The CoT traces make for fun reading—e.g. "It's not just 'a cone with eyes'; it's a whole scene with a story," or deducting points because "anthropomorphic objects are a known trope in AI art."

Related event: Multimodal LLMs Can Zero-Shot Judge Image Creativity, Study Finds(2 posts)→

Original post →

More from Research

Research channel →