CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 21 upvotes

Learning from Teacher Continuations at Student States

QUESTION — How can covariate shift and fragmented supervision be overcome during online language model distillation?

OLIVE (OnLine InterVEntion) is introduced to address limitations in existing model distillation approaches such as sequential covariate shift and fragmented supervision. In OLIVE, the evolving student generates a prefix, the teacher continues it autoregressively, and the student updates via cross-entropy on the teacher tokens. An asynchronous implementation further cuts total training time by 23.8%. Evaluated on hard reasoning and agentic tasks, OLIVE consistently outperforms existing distillation methods under matching budgets. Furthermore, continuous training with OLIVE using text from GPT-5.4-mini surpasses offline SFT from the same teacher by 13% on ScienceWorld.

An asynchronous implementation further reduces OLIVE's total training time by 23.8%.

Continuously training with OLIVE outperforms offline SFT from the same teacher by 13% on ScienceWorld.

OLIVE continues improving after offline distillation plateaus while preserving student plasticity.

shizhuo2 · 28 Sept 2026 read the original ↗
↑