Evaluation Benchmark for Keyframe Video Generation
KlingTeam · hf · 2026-07-17
This work proposes KeyFrame-Compass, designed to systematically evaluate keyframe-conditioned video generation.
Benchmark Contents
The benchmark contains 386 samples, covering:
- 3 application domains
- 2 video structure types
- 2 prompt granularities
- 2 condition formats
- 4 keyframe densities
This allows for comparing model performance differences under various settings.
Evaluation Method
The authors break down "keyframe execution" into 6 dimensions:
- Appearance
- Fidelity
- Temporal order
- Localization
- Persistence
- Uniqueness
It also evaluates overall video quality, combining evidence-driven MLLM judgments with specialized perceptual models.
Main Findings
Across 9 representative video generation systems, the authors found:
- The stricter the keyframe adherence, the more video naturalness tends to suffer
- The denser the keyframe constraints, the worse the performance
- Many open-source models even fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences
More from Multimodal
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Interactive video should be judged by responsiveness, not just frame quality — Soggy_Limit8864 · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22