CineScale: Tuning-Free Framework Lets Open-Source Video Models Generate 4K at Inference
liuziwei7 · x · 2026-09-14
Researchers from NTU and Netflix Eyeline Studios released CineScale, claimed to be the first tuning-free inference framework enabling pretrained video diffusion models to generate high-fidelity video well beyond their training resolution.
- Most video models are trained at 720p and can't exceed that at inference due to scarce 4K data and training cost.
- CineScale partitions query tokens into spatial tiles while keeping access to all global keys/values, paired with Adaptively Rectified RoPE (AR-RoPE) to preserve local geometry.
- Open-source models like Wan can now output 4K videos without any fine-tuning; the method is integrated into Wan.
- Paper includes baselines against RoPE/NTK-RoPE and cascading-step ablations.
More from Multimodal
- AI filmmaker's $10k three-week feature sparks debate over '9-dollar 7-minute video' narratives — notiansans · 2026-09-14
- GPT-6-built Human Atlas disassembles anatomy into 2,234 interactive 3D pieces — Aiden_Tech_Ai · 2026-09-14
- MiniMax 3-step Turbo LoRA tested: 2-min audio+video gen on RTX 5090 — agapes1270 · 2026-09-14
- Prompt Library: A Folder-Based Prompt Organizer Node for ComfyUI — MediocreRegister9642 · 2026-09-14
- Miroir launches Director 4K, an AI box that auto-directs multi-camera live shows — JiliJeanlouis · 2026-09-14
- Open-source ComfyUI nodes bring relight video generation via MiniMax H3 — Emotional_Example_12 · 2026-09-14