SenseNova-U1.5: an 8B encoder-free native unified multimodal model for understanding and generation
Haiwen Diao · hf · 2026-09-11
SenseNova-U1.5 is an 8B native unified multimodal model that performs visual understanding, reasoning, and generation without encoders or VAEs. It combines patch reconstruction, curated data, expert optimization, and on-policy distillation to achieve high fidelity and strong instruction following. Weights are on Hugging Face.
More from Multimodal
- Turn any article into a podcast with Meta's Muse — alexandr_wang · 2026-09-11
- GPT-6 Astra 3D Workflow: Blender MCP for Hard-Surface, TripoAI for Organic Models — majidmanzarpour · 2026-09-11
- Before Diffusion: Looking Back at VQGAN, the Pre-Diffusion Foundation of Image Generation — makeitrad1 · 2026-09-11
- ComfyUI node brings a 3-light 3D dome relighting studio to MiniMax H3 Edit — Emotional_Example_12 · 2026-09-11
- Optimized video workflow: 10s at 1MP in ~125s on an RTX 5090, with custom audio driving — Tokyo_Jab · 2026-09-11
- Seedance 2.5 generates ultra-realistic 30s Seoul lifestyle video with full prompt released — SimplyAnnisa · 2026-09-11