Diffusion Reward Models: learning the distribution of human preference instead of a single score

burny_tech · x · 2026-09-30

A new paper introduces Diffusion Reward Models (DRM), challenging a core simplification in LLM post-training: reward models collapse human preference into a single scalar score. Since different people can hold distinct, even conflicting judgments about the same response, DRM instead learns the distribution of human preference. Paper, code, and models are all open-sourced.

Related event: Tsinghua Proposes Diffusion Reward Models for Human Preferences(2 posts)→

Original post →

More from Research

Research channel →