Ostris shows DiT code-to-image reproduction progress in AR-diffusion hybrid trained on one RTX 6000 Pro

ostrisai · x · 2026-09-20

Ostris (AI Toolkit author) continues his weekend AR/diffusion hybrid image model: Qwen3-VL-4B generates image codes whose hidden states condition a frozen BFL Klein 4B diffusion decoder. He notes accurate DiT code-to-image reproduction is crucial and still has a gap to close, after only 24 hours of training on a single RTX 6000 Pro.

Related event: AI Toolkit Dev Trains AR+Diffusion Hybrid Image Model on a Single GPU in 24 Hours(3 posts)→

Original post →

More from Multimodal

Multimodal channel →