PACT: From Credit Assignment to Critic Alignment
The paper addresses token-level credit assignment in LLM reinforcement learning by formulating three regularity conditions: Completeness, Prefix Consistency, and Neutrality. Building on this unified basis, the authors introduce Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction and better align the critic. Experiments on agentic math reasoning and SWE-bench Verified show PACT substantially outperforms standard methods like GRPO and PPO.
In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively.
On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.