LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
QUESTION — How can multidimensional reward signals be constructed for legal language models?
This paper introduces LexReward, a taxonomy-driven framework for legal reward modeling. The framework characterizes legal response quality across three dimensions: Style, Element, and Chain. The authors develop rubrics for each dimension to generate pairwise preference data for DPO and reward model training. Experiments demonstrate that the learned reward models, LexRM, improve policy performance through reinforcement learning without requiring reference answers.
yidacai · 30 Sept 2026
read the original ↗