CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 111 upvotes

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

QUESTION — How to redistribute credit in code agent reinforcement learning to favor clean, targeted implementations over those containing unnecessary changes?

The paper points out that GRPO assigns identical advantages to all test-passing trajectories in code agent reinforcement learning, ignoring implementation quality. To solve this, the authors introduce GAGAR, a framework for quality-aware credit redistribution. GAGAR retains rollout groups containing both passing and failing trajectories in a shared workspace where an SFT-trained agentic grader ranks the passing candidates. It then downweights lower-ranked trajectories and rescales the advantages of passing ones to preserve their sum, resulting in more stable training and reduced length growth on MiMo-V2.6 checkpoints.

GAGAR addresses the limitation of GRPO by introducing quality-aware credit redistribution for code agents.

GAGAR performs sum-preserving redistribution of advantages to shift credit toward higher-quality implementations.

Large-scale evaluations are conducted using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters).

whatseeker · 26 Sept 2026 read the original ↗
↑