SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
This paper introduces SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling to route token-level supervision in on-policy distillation (OPD). Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. The authors also implement an engine-resident speculative verifier that preserves the trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students.
SAKI improves matched-workload rollout throughput by 4.22x using an engine-resident speculative verifier.
The method improves Mean@8 and Pass@8 across seven mathematical reasoning benchmarks for both 1.7B and 0.6B students.
Under maximal coupling, the correction probability is exactly TV(p_t, q_t).