CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 107 upvotes

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

QUESTION — What causes length inflation in on-policy distillation, and how can termination-token mismatch be resolved?

We study length inflation in on-policy distillation (OPD), identifying termination-token mismatch between base students and post-trained teachers as a primary driver of excessively long responses. Even with identical declared stopping sets, models can place stopping probabilities on different EOS tokens, suppressing the student's preferred termination. The authors show that aligning the decoding stopping set alone is insufficient, whereas treating functionally equivalent EOS tokens as a shared semantic stopping action successfully mitigates length inflation across Qwen3, Llama, and Gemma.

Termination-token mismatch between base students and post-trained teachers is an important source of length inflation.

Aligning the decoding stopping set alone is insufficient to resolve the behavior.

Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates length inflation across Qwen3, Llama, and Gemma.

kzhao5 · 17 Sept 2026 read the original ↗
↑