SenseNova U1.5-Lite release: Expert training with OPD distillation
SandyL925 · reddit · 2026-08-21
SenseTime released SenseNova U1.5-Lite, adopting a "specialized in training, unified in delivery" approach. It trains task-specific expert models for text rendering and image editing, then distills them into a single model using OPD (Online Pairwise Distillation), eliminating the need for routing or switching during inference.
Key Improvements:
- Complex Instruction Following: Consistently handles multiple constraints (subjects, counts, spatial relations, text, layout).
- Text Rendering & Dense Layouts: Supports Chinese and English in posters and infographics.
- Native 4K Generation: Maintains stable global structure.
- Native Editing: Better preservation of subject identity and geometry in multi-reference and local edits.
- Visual Understanding Boosts Generation: Transfers representations of object relations and spatial structure from understanding tasks to generation.
Benchmarks: Achieved 60.18 on Qwen-Image-Bench, 4.59 on ImgEdit, and 8.26 on GEdit-Bench-EN. The model is open-sourced on GitHub and Hugging Face.
More from Multimodal
- 3D motion graphics video created in hours using MiniMax Design — umesh_ai · 2026-08-21
- Seeking local Img2Img workflow for identity preservation on 12GB VRAM — Devilray31 · 2026-08-21
- SenseNova releases U1.5-Lite with improved text and layout control — BeCalmr · 2026-08-21
- Seedance 2.5 review: Lifelike motion and cinematic polish tested — SimplyAnnisa · 2026-08-21
- Kunlun's SkyProduction cuts AI short-drama video price to ¥0.4/sec, hits $1M revenue per show — 昆仑万维集团 · 2026-08-21
- Veo and Kling Lead: Comparison of 15 AI Video Tools — Pale_Intern5543 · 2026-08-21