CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 17 upvotes

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

QUESTION — How does on-policy self-distillation using synthetic spatial guidance improve MLLM capabilities?

The authors introduce Where-OPD, an on-policy self-distillation method for multimodal large language models that provides the teacher with textually, spatially grounded guidance identifying visual elements relevant to a query. They use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. This approach improves performance on counting, document, and chart understanding benchmarks while successfully transferring to real-world benchmarks.

Procedurally generated scenes with automatically available object identities and spatial coordinates enable scalable and annotation-free post-training.

Performance improves consistently on counting, document and chart understanding benchmarks across multiple models.

Yields a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld.

SophiaSirko · 01 Oct 2026 read the original ↗
↑