Diffusion Reward Models
QUESTION — How can reward models capture the multimodal nature of human preferences instead of relying on point estimates or fixed parametric distributions?
The authors introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over p(rmid x,y). Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, naturally representing the multimodal structure of human preferences. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger models, and improves downstream policy performance when used for RLHF training.
hbx · 27 Sept 2026
read the original ↗