CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 17 upvotes

Diffusion Reward Models

QUESTION — How can reward models capture the multimodal nature of human preferences instead of relying on point estimates or fixed parametric distributions?

The authors introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over p(rmid x,y). Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, naturally representing the multimodal structure of human preferences. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger models, and improves downstream policy performance when used for RLHF training.

hbx · 27 Sept 2026 read the original ↗
↑