LoopVL: recurrent vision-language model beats larger non-recurrent peers, shows Visual Aha Moments
RUC · hf · 2026-10-01
RUC introduces LoopVL, extending Loop Transformers to vision-language models.
- Combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules, trained from scratch via language pretraining, multimodal training, and post-training.
- Outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks.
- Observes "Visual Aha Moments": pronounced shifts in visual attention across loops.
The work offers practical evidence for recurrent vision-language modeling and a lens on how shared parameters support deeper multimodal computation.
More from Multimodal
- Supertonic TTS repo archived: development ended Sep 9, weights moved to archive namespace — JafarNajafov · 2026-10-01
- Supertonic: open-source 99M-param local TTS beats ElevenLabs in tests — JafarNajafov · 2026-10-01
- Skillry aggregates 389 viral Claude Opus 5.5 videos with copyable prompts and live remakes — CodeByPoonam · 2026-10-01
- Hands-on comparison: Boogu Image Turbo vs Krea2 Turbo vs Ideogram v4 Instant vs Fibo Lite — COMPLOGICGADH · 2026-10-01
- Stress-testing lip sync models with nine reproducible failure cases, from beards to fast speech — NewPhoneWhotiz · 2026-10-01
- Diffusers tensor parallel loading gets 2.4x faster, cuts CPU memory by 89% — RisingSayak · 2026-10-01