Tsinghua NLP Proposes Diffusion Reward Model DRM for Multimodal Human Preference

TsinghuaNLP · hf · 2026-09-29

Tsinghua NLP introduces DRM (Diffusion Reward Model), recasting reward modeling as conditional density estimation over p(r|x,y) to capture the inherently multimodal structure of human preference.

Related event: Tsinghua Proposes Diffusion Reward Models for Human Preferences(2 posts)→

Original post →

More from Research

Research channel →