CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 36 upvotes

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

QUESTION — How can multi-teacher feedback be balanced during on-policy distillation to merge specialist capabilities into a single model?

The paper reveals that traditional multi-teacher on-policy distillation (MOPD) suffers from unbalanced feedback, where instruction-following dominates student updates over mathematics. To fix this, the authors propose Domain-Normalized MOPD (DN-MOPD), which rescales each domain's feedback by its measured spread. Experiments across Qwen3.5 models and six public benchmarks show that DN-MOPD improves average scores at every size and recovers lost mathematics performance.

On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.

Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably.

Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

XINLI1997 · 28 Sept 2026 read the original ↗
↑