Rubric Rewards from Item Response Theory
The paper introduces Rubric Response Theory (RRT), which measures quality and selects criteria using item response theory when rubric criteria are monotone indicators of a shared target. Instead of adding assigned points, RRT uses a two-parameter item response model treating verdict patterns as evidence about scalar quality specific to the rubric, maximizing the local signal-to-noise ratio. A Response Parameter Network (RPN) predicts criterion difficulty and discrimination from prompt and criterion text, updated via online expectation maximization. Using Qwen3.5-4B as the policy, RRT achieves higher macro criterion scores than GRPO while reducing judge requests.
RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric.
A Response Parameter Network predicts criterion difficulty and discrimination, updated using online expectation maximization.
With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO).