CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 16 upvotes

Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

QUESTION — How can local parameter perturbations be leveraged to improve on-policy self-distillation for mathematical reasoning models?

This paper proposes Neighborhood On-Policy Self-Distillation (N-OPSD) to enhance the training of mathematical reasoning models by exploiting local parameter perturbations that reveal complementary reference-aligned corrections. Instead of a fixed parameter setting, N-OPSD builds an offline compact pool of frozen experts using greedy selection based on filtered reference-token gains, and employs online routing via MaxPeak and quantile selection. Evaluated on AIME 2024, AIME 2025, and HMMT February 2025 across Qwen3-1.7B, 4B, and 8B models, N-OPSD consistently improves the Average@12 metric over standard OPSD.

Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively.

dingyii · 30 Sept 2026 read the original ↗
↑