SenseTime open-sources 8B unified multimodal SenseNova U1.5 Lite with native 4K generation and precise editing
HeyAmit_ · x · 2026-08-21
SenseTime has open-sourced SenseNova U1.5 Lite, an 8B-parameter lightweight unified multimodal model that jointly learns visual understanding, spatial reasoning, text, generation, and editing. Highlights:
- Native 4K generation with better detail, composition, and realism
- Complex instruction following across subjects, counts, text, layouts, and styles
- Enhanced Chinese/English text rendering and multi-text layout
- Precise image editing via visual markers, bounding boxes, and multi-image references — e.g., replacing a selected infographic element while preserving its position, surrounding layout, and visual relationships
It reportedly outperforms same-size models in instruction following and edit preservation while competing with large commercial models in text rendering. The SenseNova-U1 repo has 5.1k stars on GitHub.
More from Multimodal
- Seeking local Img2Img workflow for identity preservation on 12GB VRAM — Devilray31 · 2026-08-21
- Seedance 2.5 review: Lifelike motion and cinematic polish tested — SimplyAnnisa · 2026-08-21
- SenseNova U1.5-Lite release: Expert training with OPD distillation — SandyL925 · 2026-08-21
- Kunlun's SkyProduction cuts AI short-drama video price to ¥0.4/sec, hits $1M revenue per show — 昆仑万维集团 · 2026-08-21
- Veo and Kling Lead: Comparison of 15 AI Video Tools — Pale_Intern5543 · 2026-08-21
- User Plays Pictionary Against Google Gemma 4 — tristanbob · 2026-08-21