CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 31 upvotes

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

QUESTION — How can scalar drift be eliminated in video reward models?

The paper proposes RewardVerse, a rubric-based video reward framework designed to overcome scalar drift when mapping complex video quality into a single scalar score. The framework generates explicit evaluation criteria prior to scoring, acting as a stable semantic anchor. They introduce Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm that warms up the scorer and jointly optimizes the rubric generator while aligning with human ratings. Experiments on the 16-dimensional EvalVerse benchmark demonstrate that the method mitigates scalar drift and achieves state-of-the-art performance on pointwise and pairwise evaluations.

RewardVerse introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer to mitigate scalar drift.

EddieYang428 · 19 Sept 2026 read the original ↗
↑