Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
QUESTION — How can batch-level signal imbalance be overcome in multi-reward reinforcement learning for large language models?
Multi-reward reinforcement learning optimizes multiple behavioral objectives simultaneously but often suffers from uneven learning progress due to signal imbalance. This study analyzes the phenomenon through advantage energy to show it is proportional to active-group density under GDPO normalization. To address this, the authors propose Density-Aware Reward Aggregation (DARA), which uses an inverse-square-root density correction to weight signals from less frequently active rewards, accelerating learning on tool-calling and mathematical-reasoning tasks.
DARA reaches high format compliance in up to 26% fewer training steps on tool calling.
DARA reaches near-saturated length compliance in up to 65% fewer steps on mathematical reasoning.
zhaihaotian · 30 Sept 2026
read the original ↗