SenseNova Open-Sources 8B Multimodal Model U1.5-Lite: Native 4K Generation & Precise Editing
量子位 · wechat · 2026-08-03
SenseNova has open-sourced SenseNova U1.5-Lite-Preview, a lightweight native unified multimodal model. Based on the 8B-MoT architecture and NEO-Unify, it integrates language, visual semantics, and pixel generation into a single model.
Key capabilities:
- Native 4K Output: Redesigned generation head and 4K training reduce grid artifacts and texture tearing at high resolutions.
- Complex Instructions: Handles long, multi-constraint prompts to generate high-density infographics and posters.
- Precise Editing: Supports localized edits using bounding boxes or markers (e.g., specific text replacement or lighting changes) while preserving the rest of the image.
- Multi-Reference Fusion: Can read multiple reference images simultaneously to combine layouts and content.
Benchmarks show significant improvements over the previous generation, with Qwen-Image-Bench scores rising from 47.14 to 55.20. The model is now available on GitHub, HuggingFace, and ModelScope.
More from Models
- Qwen3.8-Max initial tests show DeepSeek-level performance without degradation — gerardsans · 2026-08-03
- DeepSeek V4F Beats V4P at Chess, Writes Its Own Engine Mid-Game — teortaxesTex · 2026-08-03
- Claude Opus Falsely Claims Job Done When Context Window Fills Up — pvncher · 2026-08-03
- DeepSeek Processes 8T Tokens Daily, MiniMax Open-Sources Video Model H3 — 快鲤鱼 · 2026-08-03
- Migrating from GPT-4o to GPT-5.1: Handling RAG Agent Response Style Regressions — IncreaseLocal2574 · 2026-08-03
- Alibaba Releases Most Capable AI Model Qwen3.8-Max, Challenging OpenAI and Anthropic — The Verge AI · 2026-08-03