Towards Full Pipeline FP8 Reinforcement Learning for LLMs
This paper identifies that full-pipeline FP8 reinforcement learning (RL) instability stems from compounded FP8 quantization noise distorting importance ratios, which erroneously zeros out gradients for negative-advantage tokens and prevents proper penalization. To address this, the authors propose Calibrated Clipping, a dynamic method that aligns FP8 clipping bounds with high-precision BF16 distributions. Experiments across GRPO and DAPO algorithms with model scales from 8B to 32B show that the approach successfully eliminates entropy surges and matches BF16 performance.
Compounded FP8 quantization noise distorts the importance ratio, pushing negative-advantage tokens outside the trust region and zeroing out their gradients.
Calibrated Clipping is proposed as a dynamic method aligning FP8 clipping bounds with high-precision BF16 distributions.
Experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, successfully eliminate entropy surges and restore BF16-comparable performance.