Tsinghua NLP Proposes Diffusion Reward Model DRM for Multimodal Human Preference
TsinghuaNLP · hf · 2026-09-29
Tsinghua NLP introduces DRM (Diffusion Reward Model), recasting reward modeling as conditional density estimation over p(r|x,y) to capture the inherently multimodal structure of human preference.
- Architecture: a lightweight Diffusion Transformer, conditioned on a frozen LLM encoder, denoises Gaussian noise into a reward vector with no parametric assumption on the output distribution; one architecture handles both multi-attribute regression and pairwise preferences. At inference, N samples form an empirical reward distribution aggregable into a scalar, variance, or quantiles.
- Results: across five benchmarks, DRM matches or surpasses baselines under matched data and backbone and stays competitive with much larger discriminative, distributional, and generative RMs; it recovers multimodal reward structure that conventional heads collapse.
- Applications: uncertainty-aware rejection and LCB aggregation improve reward-model decisions; using DRM as the RLHF training-time reward improves downstream policy performance.
Related event: Tsinghua Proposes Diffusion Reward Models for Human Preferences(2 posts)→
More from Research
- Frontier AI Is a Set, Not a Point: Jagged Capabilities May Be the Steady State — vsikka · 2026-09-30
- UMAP update: new 'recursive' init scales to big data, reproducible multi-core runs — leland_mcinnes · 2026-09-30
- Researcher loses confidence in AA benchmarks, calls them "very misleading" — tianyin_xu · 2026-09-30
- Adelaide's DCSD Decouples Credit Direction and Magnitude in Self-Distillation, Beats GRPO Across 11 Benchmarks — AdelaideUniversity · 2026-09-30
- NVIDIA's HumanoidMimicGen turns one teleop demo into thousands of humanoid demonstrations — AjayMandlekar · 2026-09-30
- Color coding trick yields 2^O(sqrt(n)) depth-3 AC circuits for all symmetric Boolean functions — rrwilliams · 2026-09-30