One model streams video live and takes new instructions mid-generation

artetxem · x · 2026-10-05

artetxem demos a single model that streams video live and accepts new instructions mid-generation, adjusting output on the fly.

With reasoning off it runs super fast — 'this is what it feels like to just talk to it,' implying near-real-time conversational latency. A striking step toward interactive world-model-style video generation.

Related event: Single-Model System Streams Video in Real Time, Accepting Mid-Generation Instructions(2 posts)→

Original post →

More from Multimodal

Multimodal channel →