Unified DiT initialized from Qwen3-1.7B trains fast with 64x spatial compression

ostrisai · x · 2026-10-08

AI developer ostrisai is training a unified DiT initialized from Qwen3-1.7B, using the FLUX.2 VAE with 8x patching for 64x total spatial compression. No projection layers are used, so each block directly manipulates the latent.

Key details:

Original post →

More from Multimodal

Multimodal channel →