Training Wan 2.1 Video Model with Qwen3-VL Text Encoder
ostrisai · x · 2026-07-05
Open-source developer ostrisai experimented with adapting the video generation model Wan 2.1-1.3B to the Qwen3-VL-2B vision-language text encoder. The training was split equally at 33% each for text, VL, and a mix of both, currently limited to pre-training two linear layers for text input. The developer noted that the model adapts to the new text encoder extremely fast, having completed 25750 steps (BS=10) of training.
More from Multimodal
- Midjourney V8.2 adds personalization and shows off stylized image outputs — Mr_AllenT · 2026-07-27
- Midjourney’s image variety draws a Krea 2 comparison and asks how to reproduce it — diffusion_throwaway · 2026-07-27
- AI short film sets a 1985 dystopia to music and leans into cinema — ProfessorKey98 · 2026-07-27
- A new BOTPD episode made with Google Omni turns into an AI chase-scene parody — ScriptLurker · 2026-07-27
- A new LoRA recreates GTA: San Andreas’ classic RenderWare-era visuals — Humble-Pick7172 · 2026-07-27
- Enabling dynamic VRAM cuts LTX 2.3 video generation to 168s on an AMD R9700 — xdcfret1 · 2026-07-27