AI Toolkit author trains AR-diffusion hybrid image model in 24 hours on a single RTX 6000 Pro

ostrisai · x · 2026-09-20

AI Toolkit author Ostris hacked together an autoregressive/diffusion hybrid image model, similar to YuE2's approach: Qwen3-VL-4B generates codes for the whole image, and the hidden states of those codes are the only conditioning for a frozen BFL Klein 4B diffusion decoder. After just 24 hours on a single RTX 6000 Pro, it's learning shockingly fast.

Related event: AI Toolkit Dev Trains AR+Diffusion Hybrid Image Model on a Single GPU in 24 Hours(3 posts)→

Original post →

More from Multimodal

Multimodal channel →