Gen2Balance: text-to-video synthesis fills long-tailed data for SOTA action recognition

dimadamen · x · 2026-09-10

Gen2Balance (Univ. of Bristol & Adobe Research, ECCV 2026) tackles long-tailed video action recognition by generating synthetic clips with text-to-video models. An MLLM analyzes real exemplars to write diverse, class-faithful prompts, and a two-stage training strategy mitigates synthetic domain shift. With a released dataset of 140K generated clips across 223 classes, it beats strong baselines on UCF-LT and K100-LT, with large gains on tail and few-shot actions.

Original post →

More from Multimodal

Multimodal channel →