SenseNova Open-Sources 8B Multimodal Model for 4K Image Generation and Editing

智东西 · wechat · 2026-08-03

SenseNova recently open-sourced SenseNova U1.5-Lite-Preview, a lightweight natively unified multimodal model. With a compact scale of only 8B-MoT parameters, it integrates a full suite of capabilities including visual understanding, visual reasoning, 4K ultra-HD generation, and high-precision image editing.

Built on the self-developed NEO-Unify natively unified multimodal architecture, the model surpasses numerous larger open-source and closed-source models in both English and Chinese evaluations on the GEdit-Bench image editing benchmark. It excels at executing long, complex prompts, supports creative fusion using single or multiple reference images, and allows users to perform granular modifications to specific text, layouts, or elements via natural language or masked regions—while keeping non-edited areas completely stable. The model is now openly available on HuggingFace, GitHub, and ModelScope.

Related event: SenseTime Open-Sources 8B Multimodal Model for 4K Generation and Editing(2 posts)→

Original post →

More from Multimodal

Multimodal channel →