Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
This paper proposes Neighborhood On-Policy Self-Distillation (N-OPSD) to enhance the training of mathematical reasoning models by exploiting local parameter perturbations that reveal complementary reference-aligned corrections. Instead of a fixed parameter setting, N-OPSD builds an offline compact pool of frozen experts using greedy selection based on filtered reference-token gains, and employs online routing via MaxPeak and quantile selection. Evaluated on AIME 2024, AIME 2025, and HMMT February 2025 across Qwen3-1.7B, 4B, and 8B models, N-OPSD consistently improves the Average@12 metric over standard OPSD.
Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively.