Diffusion Reward Models: Learning Human Preference Distributions Beyond Scalar Scores
This paper introduces Diffusion Reward Models (DRM) that replace the scalar value head with a Diffusion Transformer to model full reward distributions, capturing multimodal human disagreement, enabling uncertainty-aware decisions, test-time scaling via repeated reward sampling, and improving downstream RLHF performance.
Problem: Scalar Rewards Discard Human Disagreement Structure
Standard reward models (RMs) output a single scalar for a prompt-response pair, implicitly assuming a stable "true reward." However, human annotators often disagree substantially: on HelpSteer2-Disagreements, 43.49% of helpfulness ratings span at least 2 points, 28.20% form separated clusters, and 24.34% show polarization (both low and high scores). On MultiPref, 45.64% of overall preference judgments have opposite directions, 37.23% form separated clusters. This disagreement is not mere noise but reflects genuine multimodal evaluation modes.
Method: Diffusion Transformer as Reward Head
DRM keeps a frozen LLM encoder (FsfairX-LLaMA3-RM-v0.1) and replaces the value head with a lightweight Diffusion Transformer (DiT). Given the encoder's semantic representation as condition, the DiT denoises from Gaussian noise to generate reward samples. Multiple samples yield an empirical reward distribution. Two variants are trained:
DRM-Multi-8B : 569K multi-attribute samples across 19 unified reward dimensions (missing dimensions masked).
DRM-Pref-8B : 273K pairwise preference pairs using symmetric pseudo-rewards plus a Bradley-Terry ranking objective.
Both frame reward modeling as learning the conditional reward distribution p(r|x,y).
Benchmark Results (Matched Data & Backbone)
On RewardBench v2 (average across subsets):
ArmoRM (multi-head scalar): 62.3
QRM (quantile regression): 64.1
DRM-Multi: 66.2
DRM-Pref: 65.8
DRM outperforms parametric distributional baselines without assuming Gaussian or fixed quantile families.
Does DRM Learn Human Disagreement Structure?
Distribution Matching
On HelpSteer2-Disagreements (helpfulness), DRM's sampled reward distributions are compared against repeated human ratings using Wasserstein, Jensen-Shannon, and L1 distances. DRM achieves Wasserstein 0.804 vs. empirical prior (1.030), global Gaussian (1.032), and pointwise baseline (1.032). DRM adapts its distribution per sample, not just learning a global noise level.
Multimodality Correlates with Human Disagreement
Samples are binned by human disagreement strength (low, medium, polarized). DRM's multimodality rate (fraction of multimodal sampled distributions) rises sharply:
Helpfulness: 37.6% → 55.5% → 63.2%
Correctness: 37.6% → 54.7% → 62.0%
This confirms DRM captures structured disagreement, not random variance.
Using the Distribution: Three Practical Applications
1. Selective Prediction (Uncertainty-Aware Decision Making)
For pairwise comparisons, compute the probability that B's reward samples exceed A's. Dropping the most uncertain decisions (coverage 100% → 70%) improves accuracy:
PPE Correctness (5 tasks): +2.81 points average
RMB (4 settings): +4.56 to +7.31 points
2. Best-of-N with Lower Confidence Bound (LCB)
When abstention is impossible (must pick one of N candidates), rank by mean - λ·std. LCB consistently beats mean-only ranking across candidate counts (2,4,8,16,32) on PPE and RMB (Helpfulness, Harmlessness), proving the distribution contains decision-relevant information beyond the mean.
3. Test-Time Scaling on the Reward Axis
Fix the response, repeatedly sample rewards from DRM. As sample count N grows, reward estimate stabilizes. On RewardBench v2:
DRM-Multi: 56.5 (N=1) → 65.6 (N=32)
DRM-Pref: 64.4 (N=1) → 65.7 (N=32)
This provides a second, orthogonal scaling axis to generating more responses.
Inference Cost
Default: 10 DDIM steps, 32 reward samples, batch size 1. End-to-end latency: 63.559 ms (1.63× frozen encoder baseline). Peak memory increase: +0.048 GB. Latency is dominated by DDIM steps, not sample count (32 samples add only ~1.3 ms due to parallelism). Reducing steps to 5 yields 53.755 ms (1.38×) with minor performance drop (66.2 → 65.9). Cost is negligible compared to generative RMs.
RLHF Integration
Fixed actor (Tulu3-8B-SFT), prompts (UltraFeedback), RLHF setup; only RM swapped. Results:
Arena-Hard v2: 1.3 → 2.0
MT-Bench: 73.6 → 74.8
Distributional rewards translate into policy improvements, not just RM benchmark gains.
Conclusion
DRM demonstrates that diffusion-based conditional density estimation can learn flexible reward distributions that mirror human disagreement structure. The captured uncertainty enables better decisions (selective prediction, LCB) and a novel test-time scaling dimension. While still trailing generative RMs on absolute benchmarks, DRM offers a principled path toward preserving the full information in human feedback. Code: https://github.com/thunlp/DRM, Model: https://huggingface.co/Teburile/DRM, Paper: https://huggingface.co/papers/2609.33803.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
