Study on Generalization Gap in Video Action Models

peterxichen · x · 2026-07-11

Observing the generalization capabilities of Video-Action-Models (VAMs), the author notes that while video model backbones excel at compositional generalization, VAMs often fail to achieve the same level of generalization.

They name this phenomenon the Video-Action-Generalization (VAG) gap and present a study to explain and mitigate it. The post includes a thread promising more details to follow.

Original post →

More from Research

Research channel →