Gen2Balance: text-to-video synthesis fills long-tailed data for SOTA action recognition
dimadamen · x · 2026-09-10
Gen2Balance (Univ. of Bristol & Adobe Research, ECCV 2026) tackles long-tailed video action recognition by generating synthetic clips with text-to-video models. An MLLM analyzes real exemplars to write diverse, class-faithful prompts, and a two-stage training strategy mitigates synthetic domain shift. With a released dataset of 140K generated clips across 223 classes, it beats strong baselines on UCF-LT and K100-LT, with large gains on tail and few-shot actions.
More from Multimodal
- invideo partners with Sony Liv to stream AI-generated originals — aziz4ai · 2026-09-10
- Japanese creator previews AI film 'Kaiai Village' made with Hailuo AI and Suno V6 — Hailuo_AI · 2026-09-10
- ComfyUI custom node pack banned for 'security reasons' with zero explanation from maintainers — jjjnnnxxx · 2026-09-10
- Maker chains Krea2 character sheets with MiniMax ref2vid to generate wagon scene video — StandWorth9165 · 2026-09-10
- Reusable prompt template: Watercolor Splash Fauvism style with two pure colors — LudovicCreator · 2026-09-10
- GPT6 + 3D generation builds a stunning foldable iPhone concept web demo — vista8 · 2026-09-10