Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
The authors introduce Where-OPD, an on-policy self-distillation method for multimodal large language models that provides the teacher with textually, spatially grounded guidance identifying visual elements relevant to a query. They use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. This approach improves performance on counting, document, and chart understanding benchmarks while successfully transferring to real-world benchmarks.
Procedurally generated scenes with automatically available object identities and spatial coordinates enable scalable and annotation-free post-training.
Performance improves consistently on counting, document and chart understanding benchmarks across multiple models.
Yields a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld.