CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 12 upvotes

EasyPPO: Stabilizing the Critic Is Key

QUESTION — What critic-related failure modes destabilize Proximal Policy Optimization (PPO) during large language model training?

The paper identifies the critic as a major source of instability in PPO for LLMs due to two failure modes: filtering truncated rollouts shifts the policy objective, and heterogeneous return noise lets high-variance prompts dominate updates. To address this, the authors introduce EasyPPO, incorporating actor-only overlong filtering, noise-normalized critic regression weighting loss by the inverse standard deviation of sampled returns, and moderately smaller critic mini-batches. Across FrontierCS, AIME24, and Search-R1, EasyPPO remains stable throughout training and achieves relative validation score gains of 14.89%, 2.28%, and 9.47% over vanilla PPO.

EasyPPO's best validation scores show relative gains of 14.89% on FrontierCS over PPO.

EasyPPO's best validation scores show relative gains of 2.28% on AIME24 over PPO.

EasyPPO's best validation scores show relative gains of 9.47% on Search-R1 over PPO.

wchai · 29 Sept 2026 read the original ↗
↑