CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 75 upvotes

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

QUESTION — What is the systematic failure mode known as Value Flattening in PPO critics during LLM reinforcement learning, and how can it be mitigated?

The authors uncover a systematic failure mode in PPO critics termed Value Flattening, where state values estimated via Monte Carlo change sharply while critic predictions remain flat. Analyses relate this to an implicit variance penalty in critic loss and redundant updates from temporally correlated states. To mitigate this, they introduce SParse Proximal Policy Optimization (SP^3O), which applies the value loss to only a few well-separated states in each response. Experiments on Qwen3-Base demonstrate that SP^3O with only three states supervised per response mitigates Value Flattening and consistently improves the learned policy across model sizes.

Value Flattening is a systematic failure mode where state values change sharply while critic predictions remain flat.

SP^3O applies the value loss to only a few well-separated states in each response to mitigate the issue.

Experiments on Qwen3-Base with only three states supervised per response consistently improve the learned policy across model sizes.

ramiroluo · 16 Sept 2026 read the original ↗
↑