Machine Learning Algorithms & Natural Language Processing
Oct 5, 2026 · Artificial Intelligence
Diffusion Reward Models: Learning Human Preference Distributions Beyond Scalar Scores
This paper introduces Diffusion Reward Models (DRM) that replace the scalar value head with a Diffusion Transformer to model full reward distributions, capturing multimodal human disagreement, enabling uncertainty-aware decisions, test-time scaling via repeated reward sampling, and improving downstream RLHF performance.
Diffusion ModelsDistribution ModelingHuman Preference Learning
0 likes · 20 min read
