SenseNova-U1.5: an 8B encoder-free native unified multimodal model for understanding and generation

Haiwen Diao · hf · 2026-09-11

SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs. It combines patch reconstruction, curated data, expert optimization, and on-policy distillation to achieve high fidelity and strong instruction following. Weights are on Hugging Face.

Original post →

More from Multimodal

Multimodal channel →