Abstract
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
Community
LoopVL: Recurrent Visual Intelligence
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing (2026)
- Loop Dropout: Regularizing Shared Updates in Looped Language Models (2026)
- GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning (2026)
- What Makes Recurrence Effective in Looped Language Models? (2026)
- DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models (2026)
- LoopICL: Looping a single transformer block to solve tabular tasks (2026)
- TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper