Tsinghua Proposes Diffusion Reward Models for Human Preferences
Tsinghua's NLP team introduced Diffusion Reward Models (DRM), which replace single-score reward estimation with conditional density estimation of p(r|x,y), capturing the inherently multi-modal structure of human preferences in LLM post-training.
2026-09-29 ~ 2026-09-30 · 2 related posts
- Tsinghua NLP Proposes Diffusion Reward Model DRM for Multimodal Human Preference — TsinghuaNLP · 2026-09-29
- Diffusion Reward Models: learning the distribution of human preference instead of a single score — burny_tech · 2026-09-30