Training Wan 2.1 Video Model with Qwen3-VL Text Encoder

ostrisai · x · 2026-07-05

Open-source developer ostrisai experimented with adapting the video generation model Wan 2.1-1.3B to the Qwen3-VL-2B vision-language text encoder. The training was split equally at 33% each for text, VL, and a mix of both, currently limited to pre-training two linear layers for text input. The developer noted that the model adapts to the new text encoder extremely fast, having completed 25750 steps (BS=10) of training.

Original post →

More from Multimodal

Multimodal channel →