CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 23 upvotes

PACT: From Credit Assignment to Critic Alignment

QUESTION — How can token-level credit be mathematically defined and leveraged to improve actor-critic training in LLM post-training?

The paper addresses token-level credit assignment in LLM reinforcement learning by formulating three regularity conditions: Completeness, Prefix Consistency, and Neutrality. Building on this unified basis, the authors introduce Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction and better align the critic. Experiments on agentic math reasoning and SWE-bench Verified show PACT substantially outperforms standard methods like GRPO and PPO.

In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively.

On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.

Eclipse2001 · 22 Sept 2026 read the original ↗
↑