LoopVL: recurrent vision-language model beats larger non-recurrent peers, shows Visual Aha Moments

RUC · hf · 2026-10-01

RUC introduces LoopVL, extending Loop Transformers to vision-language models.

The work offers practical evidence for recurrent vision-language modeling and a lens on how shared parameters support deeper multimodal computation.

Original post →

More from Multimodal

Multimodal channel →