Learning from Teacher Continuations at Student States
OLIVE (OnLine InterVEntion) is introduced to address limitations in existing model distillation approaches such as sequential covariate shift and fragmented supervision. In OLIVE, the evolving student generates a prefix, the teacher continues it autoregressively, and the student updates via cross-entropy on the teacher tokens. An asynchronous implementation further cuts total training time by 23.8%. Evaluated on hard reasoning and agentic tasks, OLIVE consistently outperforms existing distillation methods under matching budgets. Furthermore, continuous training with OLIVE using text from GPT-5.4-mini surpasses offline SFT from the same teacher by 13% on ScienceWorld.
An asynchronous implementation further reduces OLIVE's total training time by 23.8%.
Continuously training with OLIVE outperforms offline SFT from the same teacher by 13% on ScienceWorld.
OLIVE continues improving after offline distillation plateaus while preserving student plasticity.