Think Before You Score: Thinking Reward Model for Visual Generation
This paper introduces the Thinking Reward Model (TRM) based on a 'Think Before You Score' paradigm, which explicitly formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. To mitigate score polarization caused by conventional pairwise preference optimization, the authors propose Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing benchmarks show that TRM achieves state-of-the-art performance among open-source reward models and effectively optimizes visual generation models when used in reinforcement learning.
TRM achieves state-of-the-art performance among open-source reward models on image generation and editing benchmarks.
PD-GRPO leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring.
Using TRM as a reward for reinforcement learning consistently improves diverse visual generation models.