WanPE: 397B Prompt Enhancer Boosts 30s Video Preference by 50.86 Points
Yubo Zhu · hf · 2026-09-25
WanPE is a 397B-parameter prompt enhancement model trained on 1.05M real-world videos for director-level cinematic planning in text-to-video generation.
- Method: Shot-level cinematic plans via video-grounded reverse construction, with Semantic-Consistency GRPO (SC-GRPO) preserving user requirements across shots and time
- Evaluation: Accompanied by WanPEval, a human-annotated testbed covering 5–30s durations, supported by 11K blind pairwise assessments
- Results: Powering Wan3.0's generator, human preference rises 10.66–18.84 points at 5–15s and a dramatic 50.86 points at 30s; leads all evaluated commercial offerings at 5–15s and stays competitive with Seedance 2.5 at 30s
- Ablations: Reverse construction clearly beats forward rewriting; SC-GRPO robustly preserves semantic fidelity across scales
More from Multimodal
- User generates a music video from old material with Opus 5.5 — repligate · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25
- One prompt: Claude agent wired to Runway MCP delivers a Netflix-style superintelligence doc — CurieuxExplorer · 2026-09-25
- Seedance 2.5 + GPT Image 2.5 One-Shot the Most Iconic Sci-Fi Rivalry — CurieuxExplorer · 2026-09-25
- Opus 5.5 writes songs and music videos entirely from code in a playable demo — pbaylies · 2026-09-25
- Study: Reasoning hurts 15.7% of multimodal embeddings; training-free SURE router fixes it — _reachsumit · 2026-09-25