UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
This paper introduces UniEvo-VL, a self-evolving framework for multimodal models to learn from constructive self-correction feedback during test-time compute. Instead of relying on a separate larger teacher, UniEvo-VL leverages the model's own self-critiques as privileged information, with a single multimodal model acting as both teacher and student under different contexts. Training minimizes the per-state divergence between their denoising diffusion distributions over the student's sampling trajectories. Experiments built on Qwen-image-2512 show significant performance gains from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA.
UniEvo-VL improves the image generation capabilities of multimodal models without relying on an external teacher.
Performance on GenEval increases from 0.747 to 0.808.
Performance on GenEval2 Soft-TIFA increases from 32.97 to 35.53.