SenseNova-U1.5 Technical Report: VAE-Free Native 4K Image Generation

Secret_Yak2496 · reddit · 2026-09-16

SenseNova released the technical report for SenseNova-U1.5, an 8B native unified model for image understanding, generation and editing that skips external visual encoders and VAEs — images map directly to visual tokens, each covering a 32×32-pixel region.

Key methods: spatially joint reconstruction via Pixel Shuffle + 3×3 convs (fixing U1's patch-boundary artifacts from independent token decoding), resolution-aware noise conditioning extended to 4096×4096, and a 'specialize, then unify' scheme — four RL experts (aesthetics, bilingual text rendering, infographics, editing) consolidated via multi-expert on-policy distillation. Training added 59M text-image pairs from 78 sources and 38M editing examples, with 88.2% of effective generation volume above 1024². Known issues remain with dense small text, small faces/hands and multi-reference drift. Open-sourced on GitHub and Hugging Face.

Original post →

More from Multimodal

Multimodal channel →