SenseNova-U1.5: 8B encoder-free unified model does visual understanding and generation in one

KyeGomezB · x · 2026-09-14

SenseTime's SenseNova-U1.5 is an 8B-MoT natively unified multimodal model with an encoder-free, VAE-free architecture that understands, reasons about, generates, and edits images directly in pixel space, including native 4K generation.

Key details:

The paper argues visual understanding and generation can share one native representation rather than separate systems.

Original post →

More from Multimodal

Multimodal channel →