Training Qwen3-VL-4B into an AR image generator with Cosmos tokenizer
ostrisai · x · 2026-10-01
ostrisai shares an in-progress experiment training Qwen3-VL-4B-Instruct as an autoregressive image generator, using nvidia/Cosmos-0.1-Tokenizer-DI16x16 as the image tokenizer. At 256 resolution the model already produces images from default AI Toolkit prompts, demonstrating the approach works at an early stage.
More from Multimodal
- One creator's short-film workflow with Nano Banana Pro + h3, and whether LTX 2.5 can match it — No_Reference_7678 · 2026-10-01
- Ideogram 4.5 Surfaces in New Teaser for Text-to-Image Model — aziz4ai · 2026-10-01
- MiniMax-H3 Is Secretly a Strong Image Generator in ComfyUI — solomars3 · 2026-10-01
- LEAP uses learned block-wise evidence retrieval to boost hour-scale audio-video QA by up to 16.8% — Juyi Lin · 2026-10-01
- Video Genie brings natural-language AI video editing to Mac with 1,000+ community skills — realmeetjames · 2026-10-01
- You can animate 3D Gaussian splats — PlayCanvas demos an animated splat dino — willeastcott · 2026-10-01