Real-time generation demo shows hybrid image model generating whole-image codes first

ostrisai · x · 2026-09-20

Ostris shares a real-time generation clip of his AR/diffusion hybrid image model (Qwen3-VL-4B + frozen BFL Klein 4B): codes for the whole image are generated first, then decoded — with no inference optimizations applied.

Original post →

More from Multimodal

Multimodal channel →