Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
This work introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically to improve LLM reasoning through reinforcement learning. By coupling normalized policy entropy with response advantage signs, EAPO reinforces surprising successes at high-entropy positions while penalizing confident failures at low-entropy spots, all while attenuating penalties at uncertain locations to preserve recovery opportunities. Derived directly from existing rollout signals without extra supervision, EAPO promotes effective problem exploration and broad candidate generation across various reasoning backbones.
EAPO treats success and failure asymmetrically by coupling normalized policy entropy with the sign of the response advantage.
The method derives token-level credit directly from existing rollout signals without additional supervision.
EAPO achieves the best overall performance across a range of reasoning tasks on both base and reasoning backbones.