Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
L'article révèle que la distillation en politique multi-enseignants (MOPD) traditionnelle souffre de retours déséquilibrés, où le suivi des instructions domine les mises à jour au détriment des mathématiques. Les auteurs proposent Domain-Normalized MOPD (DN-MOPD), qui redimensionne les retours de chaque domaine en fonction de leur dispersion mesurée. Les expériences sur les modèles Qwen3.5 et six benchmarks publics montrent que DN-MOPD améliore les scores moyens à toutes les tailles.
On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.
Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably.
Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.