Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
QUESTION — How can native reflection and image generation capabilities be jointly trained in a unified multimodal model using reinforcement learning?
The authors introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection and image repair loops within a unified multimodal model. Sibling trajectories share an initial image to compare reflection strategies via group-relative advantage, updating both reflection tokens and flow-based revisions without needing an external inference verifier. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, with transfer gains on WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63).
On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Ziqi · 28 Sept 2026
read the original ↗