Ostris shows DiT code-to-image reproduction progress in AR-diffusion hybrid trained on one RTX 6000 Pro
ostrisai · x · 2026-09-20
Ostris (AI Toolkit author) continues his weekend AR/diffusion hybrid image model: Qwen3-VL-4B generates image codes whose hidden states condition a frozen BFL Klein 4B diffusion decoder. He notes accurate DiT code-to-image reproduction is crucial and still has a gap to close, after only 24 hours of training on a single RTX 6000 Pro.
More from Multimodal
- Halloween AI video demo with admittedly rough lip sync — DavidmComfort · 2026-09-20
- Sufi Qawwali and Tabla-style music LoRAs released for AI music generation — -becausereasons- · 2026-09-20
- Hailed as the best fully AI-generated short film so far, by javilopen — techhalla · 2026-09-20
- One-Person AI-Generated Series 'Guixu' Hits 9 Episodes at ~$20K Total Cost — dotey · 2026-09-20
- Woken by his cat, this creator shipped "Kittens on Ukulele Island" — bennash · 2026-09-20
- AI-Generated Music Video Is Flawless in Every Way Except the Genuinely Awful Singing — aronchick · 2026-09-20