FrameMorrow Selects History Frames by Predicted Future Needs, Boosting 11 Long-Video Generators
NationalUniversityofSingapore · hf · 2026-10-05
NTU Singapore proposes FrameMorrow for long-horizon video generation: existing methods judge historical relevance from current content, but seemingly irrelevant history may matter later.
Method: predict a small set of prospective tokens representing future information needs, and use them to select relevant explicit history frames—rather than model-internal states—enabling plug-and-play integration even with closed-source generators at little extra inference cost.
Results: Evaluated across 5 benchmarks and 11 generators (long-video, interactive generation, action-conditioned world models), consistently improving long-range consistency, visual quality, and action alignment.
More from Multimodal
- One Prompt Gets Fable to Generate a 5,000-Year History of China Video — FuSheng_0306 · 2026-10-05
- UniMate Releases 3D Rigged Skeleton Models on Hugging Face, ComfyUI Integration Proposed — RazsterOxzine · 2026-10-05
- Full Making-of Released for AI Music Video 'Close All the Windows' — All Prompts and Orchestration Logs Included — Afinetheorem · 2026-10-05
- OctLLM encodes 3D geometry as explicit octree token sequences without sacrificing language ability — _akhaliq · 2026-10-05
- Gemma 31b 'Grand Horror' + H3 generates atmospheric horror video — jrexthrilla · 2026-10-05
- EditHero: First benchmark for long-horizon part-level 3D editing compares agentic vs non-agentic approaches — _akhaliq · 2026-10-05