Local Vision Aids Generalization

qualcomm · hf · 2026-07-18

This work discusses **locality and length generalization in vision models**. Human vision relies on a sequence of local, foveated observations, whereas most computer vision models process inputs globally all at once. The authors ask: beyond aligning better with biological vision, do these "step-by-step local perception" vision models offer genuine computational advantages. Through experiments on simple visual tasks requiring the aggregation of local information across images, they found that many vision models, similar to language models, learn "global shortcuts" and thus fail to generalize when task length or complexity increases. Furthermore, a **recursive vision strategy with strictly local perception** can mitigate this issue, enabling models to generalize more robustly on these tasks. The authors conclude that **local attention may be an overlooked, critical component for robust combinatorial generalization**.

Original post →

More from Research

Research channel →