Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
The paper reveals that traditional multi-teacher on-policy distillation (MOPD) suffers from unbalanced feedback, where instruction-following dominates student updates over mathematics. To fix this, the authors propose Domain-Normalized MOPD (DN-MOPD), which rescales each domain's feedback by its measured spread. Experiments across Qwen3.5 models and six public benchmarks show that DN-MOPD improves average scores at every size and recovers lost mathematics performance.
On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.
Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably.
Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.