OmniVBench: a 12k-checklist benchmark and 340K-sample dataset for omni reference-to-video generation
Wenxue Li · hf · 2026-09-21
OmniVBench and the Omni-R2V Dataset target evaluation and training gaps in emerging "omni" reference-to-video (R2V) generation.
Benchmark: Spans 7 task families and 18 fine-grained tasks across content, motion, style, structure, narrative, and multi-reference settings. It introduces factor-grounded evaluation with 12,172 case-specific checklist items checking whether reference factors are preserved, correctly disentangled and bound to targets, and properly realized.
Dataset: Built from large-scale professional footage, Omni-R2V offers 340K processed training samples with scalable pipelines for reference-target pair construction.
Findings: Evaluations of advanced open- and closed-source R2V models reveal clear performance gaps across task families and dimensions.
More from Multimodal
- Uncensored Qwen-Image-2.1 GGUF quantization trends on Hugging Face — abenzerps · 2026-09-21
- Custom CUDA shim runs Stable Diffusion on Mac faster than RTX 5090 on Windows, up to 61% quicker — LioDavinchy · 2026-09-21
- Qwen-Image-2.1 hands-on: trails Krea2 in T2I quality and Flux2Klein in editing flexibility — Capitan01R- · 2026-09-21
- Qwen 2.1 T2I & edit test: Reddit calls it the open-source Nano Banana moment — LongjumpingGur7623 · 2026-09-21
- GPT-Image 2.5 proves surprisingly good at pixel art, animated via Seedance 2.5 — Aiden_Tech_Ai · 2026-09-21
- Qwen-Image 2.1 runs on SGLang: image generation in 18.7s on a single RTX 4090 — Alibaba_Qwen · 2026-09-21