SenseTime's SenseNova-U1.5 hits Hugging Face: 8B encoder-free multimodal model with native 4K

liuziwei7 · x · 2026-09-11

SenseTime's SenseNova-U1.5 is now on Hugging Face: a native unified 8B-MoT multimodal model that handles understanding, reasoning, and visual generation in one model.

Notably, it drops the visual encoder and VAE entirely, scaling to native 4K resolution. Weights are publicly available for download.

Related event: SenseTime Releases SenseNova-U1.5, an 8B Encoder-Free Unified Multimodal Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →