Sci-VBench: Evaluating Knowledge-Intensive Video Generation in Science
Diandian Zhang · hf · 2026-08-11
Sci-VBench is a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains.
- Dataset: Contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Requires models to generate temporally rich videos grounded in scientific reasoning.
- Protocol: Establishes a rubric-based evaluation protocol where both non-expert human evaluators and MLLM-as-Judge systems achieve high agreement with experts, supporting reproducible evaluation at scale.
- Findings: Benchmarks 16 frontier proprietary and open-source models. While automatic perceptual-quality scores cluster tightly, performance on Prompt Grounding and Scientific/Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. Visual realism advances have not yet translated into reliable modeling of scientific dynamics.
More from Multimodal
- AI Video Fail: Recreating the Iconic Miami Vice Scene with Bert and Ernie — dreamwieber · 2026-08-11
- Grok Image 2.0 Tested: Surgical Editing and Sharp Text Rendering — minchoi · 2026-08-11
- 10-Minute AI-Generated Film Goes Viral with 300K Views in a Day — ZabihullahAtal · 2026-08-11
- MiniMax M3 Tested: Generates High-Quality Video in Under 5 Minutes — Fear_ltself · 2026-08-11
- RefCaptioner: Precise Multi-Reference Image Grounding in Video Captioning — 机器之心 · 2026-08-11
- One Stylus Tap Turns Raw Ingredients Into Gourmet Ramen Using AI Video — SimplyAnnisa · 2026-08-11