CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients for multi-reward learning in Group Relative Policy Optimization. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization, preventing large-scale reward components from dominating. We evaluate CorrGRPO on code generation, tool calling, and agent security using models ranging from 0.5B to 8B parameters. Results show improvements across all three domains.
CorrGRPO converts pairwise covariances into Pearson correlation coefficients to balance the influence of differently scaled rewards.
The method is evaluated on code generation, tool calling, and agent security using models ranging from 0.5B to 8B parameters.