Unified DiT initialized from Qwen3-1.7B trains fast with 64x spatial compression
ostrisai · x · 2026-10-08
AI developer ostrisai is training a unified DiT initialized from Qwen3-1.7B, using the FLUX.2 VAE with 8x patching for 64x total spatial compression. No projection layers are used, so each block directly manipulates the latent.
Key details:
- X0 prediction plus alternating padding/slicing blocks effectively prevents patch artifacts
- The trick only works without projection layers, so latent channels must match hidden state size
- Training is fast; sample snapshots after 2 days look promising. A second run initialized from a 4B-pruned Hidream-01 with Wan 2.1 VAE (patch 16) targets 128x spatial and 4x temporal compression
More from Multimodal
- ElevenLabs Brings 11 Speech Models to OpenRouter With 50% Off Until October 19 — lukeharries · 2026-10-08
- Nano Banana 2.1 spotted in Google Flow: one-prompt movie collages from end to start — socialwithaayan · 2026-10-08
- Fresh batch of Nano Banana 2.1 generated images leaks on X — socialwithaayan · 2026-10-08
- Nano Banana 2.1 is one day old and its images are already everywhere — socialwithaayan · 2026-10-08
- TryLiveAvatar launches 30 full-body live avatars with 6 emotions, free to use — aftahi_ai · 2026-10-08
- Dev tests AI-generated games: models are good — at unoriginal ideas — zeeg · 2026-10-08