CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 13 upvotes

Rubric Rewards from Item Response Theory

QUESTION — How does Rubric Response Theory (RRT) solve the rubric reward aggregation problem in reinforcement learning?

The paper introduces Rubric Response Theory (RRT), which measures quality and selects criteria using item response theory when rubric criteria are monotone indicators of a shared target. Instead of adding assigned points, RRT uses a two-parameter item response model treating verdict patterns as evidence about scalar quality specific to the rubric, maximizing the local signal-to-noise ratio. A Response Parameter Network (RPN) predicts criterion difficulty and discrimination from prompt and criterion text, updated via online expectation maximization. Using Qwen3.5-4B as the policy, RRT achieves higher macro criterion scores than GRPO while reducing judge requests.

RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric.

A Response Parameter Network predicts criterion difficulty and discrimination, updated using online expectation maximization.

With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO).

Miladyz · 28 Sept 2026 read the original ↗
↑