Multimodal Models Can Act as Zero-Shot Reward Models

_akhaliq · x · 2026-07-15

A paper titled "Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation" reveals that pretrained multimodal large language models (MLLMs) can act as zero-shot reward models for text-to-image generation. They can be used to evaluate whether the generated results accurately match the text descriptions.

Related event: ByteDance Introduces SpectraReward: Zero-Shot MLLM as Image Reward Model(3 posts)→

Original post →

More from Multimodal

Multimodal channel →