SenseNova-U1.5: 8B encoder-free unified model does visual understanding and generation in one
KyeGomezB · x · 2026-09-14
SenseTime's SenseNova-U1.5 is an 8B-MoT natively unified multimodal model with an encoder-free, VAE-free architecture that understands, reasons about, generates, and edits images directly in pixel space, including native 4K generation.
Key details:
- Strengthens its visual interface via spatially coherent patch reconstruction, scaled with curated generation/editing data, improved task formulation, and structural prompt enhancement
- Post-training optimizes specialist experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, consolidated through multi-expert on-policy distillation
- Evaluations show major gains in image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation while preserving instruction following, subject identity, and unmodified regions
The paper argues visual understanding and generation can share one native representation rather than separate systems.
More from Multimodal
- Marigold V2 repurposes diffusion Transformers for dense prediction, trainable on a single GPU in days — RexDouglass · 2026-09-14
- Rigged Mesh Can Still Fail in Motion: Lessons from Three Chibi 3D Character Tests — softmarshmallow · 2026-09-14
- Comic-book heroes recreated with 'GPT-6 Astra', fal camera control, Three.js — OdinLovis · 2026-09-14
- AI-generated anime battle: Crazy Rari vs Weeping Argent showcase video — abokalypsis · 2026-09-14
- Google AI reconstructs an elderly couple's unrecorded first meeting — at what cost to memory? — VraserX · 2026-09-14
- Running local i2i/i2v on a 3060 12GB: Wan2.2 and MiniMax-H3 fall short, alternatives wanted — Legal-Respond-6725 · 2026-09-14