Prototype video model uses Wan 2.1 1.3B with 16x spatial and 8x temporal compression
ostrisai · x · 2026-07-27
The author is experimenting with a pixel-space video model built from Wan 2.1 1.3B as the global blocks, plus modified hourglass transformer blocks on both ends.
They say they are pretraining the hourglass part first. The global blocks use an equivalent of 16 spatial and 8 temporal compression, suggesting a heavily compressed video representation pipeline.
More from Multimodal
- GPT-5.6 Sol generates believable low-poly 3D scenes with minimal prompting — Dimillian · 2026-07-27
- LTX 2.3 close-up video looks sharper at 1080p, but a 6-second clip takes 12–15 minutes on an RTX 3060 — iiTzMYUNG · 2026-07-27
- Opus 5 one-shot a Paper Mario–style game prototype in a single run — Acid_God_ · 2026-07-27
- Runway demo turns a snowfall dance into a cinematic magical-realist video — LudovicCreator · 2026-07-27
- A Midjourney origami prompt turns any subject into a silhouette test — tisch_eins · 2026-07-27
- Midjourney 8.2 is said to outperform Leonardo AI’s Pro Upscaler — aziz4ai · 2026-07-27