CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 80 upvotes

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

QUESTION — How can multimodal models be improved through test-time self-correction feedback without requiring an external teacher?

This paper introduces UniEvo-VL, a self-evolving framework for multimodal models to learn from constructive self-correction feedback during test-time compute. Instead of relying on a separate larger teacher, UniEvo-VL leverages the model's own self-critiques as privileged information, with a single multimodal model acting as both teacher and student under different contexts. Training minimizes the per-state divergence between their denoising diffusion distributions over the student's sampling trajectories. Experiments built on Qwen-image-2512 show significant performance gains from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA.

UniEvo-VL improves the image generation capabilities of multimodal models without relying on an external teacher.

Performance on GenEval increases from 0.747 to 0.808.

Performance on GenEval2 Soft-TIFA increases from 32.97 to 35.53.

fangwu97 · 30 Sept 2026 read the original ↗
↑