Prototype video model uses Wan 2.1 1.3B with 16x spatial and 8x temporal compression

ostrisai · x · 2026-07-27

The author is experimenting with a pixel-space video model built from Wan 2.1 1.3B as the global blocks, plus modified hourglass transformer blocks on both ends.

They say they are pretraining the hourglass part first. The global blocks use an equivalent of 16 spatial and 8 temporal compression, suggesting a heavily compressed video representation pipeline.

Original post →

More from Multimodal

Multimodal channel →