When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), identifying termination-token mismatch between base students and post-trained teachers as a primary driver of excessively long responses. Even with identical declared stopping sets, models can place stopping probabilities on different EOS tokens, suppressing the student's preferred termination. The authors show that aligning the decoding stopping set alone is insufficient, whereas treating functionally equivalent EOS tokens as a shared semantic stopping action successfully mitigates length inflation across Qwen3, Llama, and Gemma.
Termination-token mismatch between base students and post-trained teachers is an important source of length inflation.
Aligning the decoding stopping set alone is insufficient to resolve the behavior.
Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates length inflation across Qwen3, Llama, and Gemma.