LoopVL: Recurrent Visual Intelligence
QUESTION — Can Loop Transformers be effectively extended to vision-language models for recurrent computation?
The authors introduce LoopVL to study the extension of Loop Transformers to vision-language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. Trained from scratch through language pre-training, multimodal training, and post-training, LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. Additionally, the model exhibits Visual Aha Moments characterized by pronounced shifts in visual attention across loops.
Gtime666 · 29 Sept 2026
read the original ↗