AI Toolkit author trains AR-diffusion hybrid image model in 24 hours on a single RTX 6000 Pro
ostrisai · x · 2026-09-20
AI Toolkit author Ostris hacked together an autoregressive/diffusion hybrid image model, similar to YuE2's approach: Qwen3-VL-4B generates codes for the whole image, and the hidden states of those codes are the only conditioning for a frozen BFL Klein 4B diffusion decoder. After just 24 hours on a single RTX 6000 Pro, it's learning shockingly fast.
More from Multimodal
- Halloween AI video demo with admittedly rough lip sync — DavidmComfort · 2026-09-20
- Sufi Qawwali and Tabla-style music LoRAs released for AI music generation — -becausereasons- · 2026-09-20
- Hailed as the best fully AI-generated short film so far, by javilopen — techhalla · 2026-09-20
- One-Person AI-Generated Series 'Guixu' Hits 9 Episodes at ~$20K Total Cost — dotey · 2026-09-20
- Woken by his cat, this creator shipped "Kittens on Ukulele Island" — bennash · 2026-09-20
- AI-Generated Music Video Is Flawless in Every Way Except the Genuinely Awful Singing — aronchick · 2026-09-20