NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling
NationalUniversityofSingapore · hf · 2026-09-30
Researchers at the National University of Singapore introduce StoryEngine, a state-grounded agentic framework for long-form video storytelling.
Problem: Existing agentic multi-shot video generation relies on textual shot plans or generated pixels, with no explicit mechanism to propagate event consequences or maintain world state across shots—causing inaccurate reconstruction of visual details and visual drift that erodes narrative coherence.
Approach:
- Separates authoritative semantic plans from unreliable visual observations
- Maintains structured representations of entity placement and story-relevant states, propagating event-induced changes into per-shot start/end states
- Builds canonical references for recurring entities/environments and compiles state and visual constraints into executable render plans
- Uses a bounded evaluation-guided repair loop to fix local state inconsistencies before they propagate
The team also releases a benchmark spanning diverse scenarios and visual styles, measuring storytelling quality, narrative coherence, and visual consistency. StoryEngine consistently outperforms state-of-the-art methods across all dimensions.
More from Multimodal
- Four imaginary tokens for Midjourney v8.2 produce memory ghosts and bone echoes — LudovicCreator · 2026-09-30
- Hyper-personalized music is BS: music is culture and inherently social, argues developer — jordiponsdotme · 2026-09-30
- One Year of Local Image Generation: Why Civitai and ComfyUI Both Fall Short — BenDLH · 2026-09-30
- Opus made a launch video for Violetto 1B in 50 minutes amid zero media coverage — tensorqt · 2026-09-30
- Meshy hits $100M ARR in under two years as GPT-6 Astra stirs the AI 3D debate — 量子位 · 2026-09-30
- LoRA adapters break on distilled video models, long post explains why — burkov · 2026-09-30