Smaller Models, Better Rejects: Preference Distillation Scaling
This research investigates preference distillation and challenges the assumption that rejects must come from student-scale or self-generated responses. The authors discover that smaller frozen models can generate rejects using less inference compute while training stronger students on code generation and mathematical reasoning. Supported by a derived finite-horizon utility bound, they propose interventions such as mixing rejects from smaller and student-scale models, shuffling code tokens while retaining task structure, and selecting candidates with lower likelihood under the reference policy to enhance net transfer.
Smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning.
Mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases.
Lower-likelihood selections outperform higher-likelihood ones for every source.