Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
This study investigates specialist models trained exclusively on raw question-answer pairs. By leveraging student distillation as an agnostic probe, the authors discover that specialist optimization implicitly selects from a latent trajectory space. Empirical analysis across 27 specialist-student pairings reveals that their specialization and generalization profiles correlate exceptionally strongly. Furthermore, explicitly controlling the specialist's distributional drift allows systematic adjustment of the trade-off between domain precision and general capability across both teacher and distilled student models.
The study analyzes across 27 specialist--student pairings.
Specialization--generalization profiles between specialists and students correlate exceptionally strongly.