OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation

Wenxue Li · hf · 2026-09-21

OmniVBench and the Omni-R2V Dataset target evaluation and training gaps in emerging "omni" reference-to-video (R2V) generation.

Benchmark: Spans 7 task families and 18 fine-grained tasks across content, motion, style, structure, narrative, and multi-reference settings. It introduces factor-grounded evaluation with 12,172 case-specific checklist items checking whether reference factors are preserved, correctly disentangled and bound to targets, and properly realized.

Dataset: Built from large-scale professional footage, Omni-R2V offers 340K processed training samples with scalable pipelines for reference-target pair construction.

Findings: Evaluations of advanced open- and closed-source R2V models reveal clear performance gaps across task families and dimensions.

Original post →

More from Multimodal

Multimodal channel →