Bellman Policy Optimization
QUESTION — How can the reasoning capabilities of large language models be improved via reinforcement learning with verifiable rewards without estimating intermediate state values?
The paper introduces Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD) for reinforcement learning with verifiable rewards (RLVR). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective, avoiding the need to estimate intermediate state values. The authors prove it shares the same unique optimal solution as the original PMD objective and derive a practical loss function whose mismatch-correction weight uses a smoothed ratio of complementary token probabilities.
Mat3579 · 14 Sept 2026
read the original ↗