Training Wan 2.1 Video Model with Qwen3-VL Text Encoder
ostrisai · x · 2026-07-05
Open-source developer ostrisai experimented with adapting the video generation model Wan 2.1-1.3B to the Qwen3-VL-2B vision-language text encoder. The training was split equally at 33% each for text, VL, and a mix of both, currently limited to pre-training two linear layers for text input. The developer noted that the model adapts to the new text encoder extremely fast, having completed 25750 steps (BS=10) of training.
More from Multimodal
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11