Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
The paper points out that GRPO assigns identical advantages to all test-passing trajectories in code agent reinforcement learning, ignoring implementation quality. To solve this, the authors introduce GAGAR, a framework for quality-aware credit redistribution. GAGAR retains rollout groups containing both passing and failing trajectories in a shared workspace where an SFT-trained agentic grader ranks the passing candidates. It then downweights lower-ranked trajectories and rescales the advantages of passing ones to preserve their sum, resulting in more stable training and reduced length growth on MiMo-V2.6 checkpoints.
GAGAR addresses the limitation of GRPO by introducing quality-aware credit redistribution for code agents.
GAGAR performs sum-preserving redistribution of advantages to shift credit toward higher-quality implementations.
Large-scale evaluations are conducted using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters).