Diffusion Reward Models: learning the distribution of human preference instead of a single score
burny_tech · x · 2026-09-30
A new paper introduces Diffusion Reward Models (DRM), challenging a core simplification in LLM post-training: reward models collapse human preference into a single scalar score. Since different people can hold distinct, even conflicting judgments about the same response, DRM instead learns the distribution of human preference. Paper, code, and models are all open-sourced.
Related event: Tsinghua Proposes Diffusion Reward Models for Human Preferences(2 posts)→
More from Research
- GamowLabs' LabBench: AI Agents Struggle to Pick the Next Best Wet-Lab Experiment — danielmckinn0n · 2026-09-30
- Toby Ord: Opus 5.5 system card yields lambda values matching OpenAI's agent parallelism estimates — tobyordoxford · 2026-09-30
- davidad Ships Interactive Playground to Test His Bayesian Cognition Model Across Systems — davidad · 2026-09-30
- Microsoft Research's Social-R1 uses RL to train genuine social reasoning in AI — burkov · 2026-09-30
- Verifying all of nixpkgs: 95% of machine code for ~$134M by end of 2027, study models — ctjlewis · 2026-09-30
- 'LLMs as a Cognitive Virus' paper with Krakauer and Levin questioned for weak LLM specificity — hn1000 · 2026-09-30