CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 31 upvotes

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

QUESTION — How can large-scale rewards be prevented from dominating normalization and suppressing signals from smaller-scale rewards in multi-reward Group Relative Policy Optimization?

We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients for multi-reward learning in Group Relative Policy Optimization. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization, preventing large-scale reward components from dominating. We evaluate CorrGRPO on code generation, tool calling, and agent security using models ranging from 0.5B to 8B parameters. Results show improvements across all three domains.

CorrGRPO converts pairwise covariances into Pearson correlation coefficients to balance the influence of differently scaled rewards.

The method is evaluated on code generation, tool calling, and agent security using models ranging from 0.5B to 8B parameters.

hubin · 29 Sept 2026 read the original ↗
↑