First samples from Ostris's AR-diffusion hybrid show Qwen3-VL-4B grasping prompt-to-image generation

ostrisai · x · 2026-09-20

Ostris shares first samples from his AR/diffusion hybrid image model (Qwen3-VL-4B autoregressive model + frozen BFL Klein 4B diffusion decoder), generated with AI Toolkit's default prompts. The AR model has clearly learned to produce images matching the prompts.

Related event: AI Toolkit Dev Trains AR+Diffusion Hybrid Image Model on a Single GPU in 24 Hours(3 posts)→

Original post →

More from Multimodal

Multimodal channel →