SenseTime releases open-source SenseNova U1.5 with MoT architecture

multimodalart · x · 2026-08-21

SenseTime released the open-source image generation model SenseNova-U1.5-8B-MoT under the Apache 2.0 license. It uses a Mixture of Transformers (MoT) architecture, eliminating the need for VAEs, text encoders, or DiTs. Text and image tokens use different weights and attend to each other, denoising in pixel space. Benchmarks show it matches Nano Banana 2, and a demo is now available.

Original post →

More from Multimodal

Multimodal channel →