Tsinghua Proposes Diffusion Reward Models for Human Preferences

Tsinghua's NLP team introduced Diffusion Reward Models (DRM), which replace single-score reward estimation with conditional density estimation of p(r|x,y), capturing the inherently multi-modal structure of human preferences in LLM post-training.

2026-09-29 ~ 2026-09-30 · 2 related posts