SenseTime open-sources SenseNova U1.5 Lite: 8B unified multimodal model with native 4K generation
aftahi_ai · x · 2026-08-21
SenseTime released SenseNova U1.5-8B-MoT, an open-source lightweight unified multimodal model covering visual understanding, generation, and editing, built for real-world visual workflows:
- Native 4K generation with better detail, composition, and realism
- Stronger complex instruction following across subjects, counts, text, layouts, and styles
- Improved Chinese/English text rendering and multi-text layout
- Precise image editing via visual markers, bounding boxes, and multi-image references
At 8B parameters it reportedly outperforms same-size models in instruction following and editing preservation while competing with large commercial models on text rendering. An 8-step LoRA variant for faster inference was also released. The GitHub repo (5.1k stars) includes ComfyUI integration, training, and evaluation code.
More from Models
- Speculation: Google to announce new Gemma model at SF event — Porespellar · 2026-08-21
- DeepSeek Flash usage up 10-100x due to ultra-cheap inference — bindureddy · 2026-08-21
- DeepSeek V4 tested with poor aesthetics, suspected hidden CoTs — teortaxesTex · 2026-08-21
- Gemini 3.7 Flash scores 84.6% on ARC-AGI-2 at $0.25 per task — typewriters · 2026-08-21
- Cisco Open Sources Antares Security Model, 3B Matches GPT-5.5 — aminkarbasi · 2026-08-21
- Claim: Anthropic Broke Opus with Text Watermarks — nikvassev · 2026-08-21