MLLMs Serve as Zero-Shot Reward Models for Text-to-Image

_akhaliq · x · 2026-07-15

This paper proposes that pretrained multimodal large language models (MLLMs) can directly serve as zero-shot reward models for text-to-image generation, evaluating the quality of the generated outputs.

Essentially, without requiring specific training for reward scoring, the models can leverage their existing vision-language capabilities to score or rank generated images, thereby helping to optimize text-to-image systems.

Related event: ByteDance Introduces SpectraReward: Zero-Shot MLLM as Image Reward Model(3 posts)→

Original post →

More from Multimodal

Multimodal channel →