SenseNova-U1.5 Technical Report: VAE-Free Native 4K Image Generation
Secret_Yak2496 · reddit · 2026-09-16
SenseNova released the technical report for SenseNova-U1.5, an 8B native unified model for image understanding, generation and editing that skips external visual encoders and VAEs — images map directly to visual tokens, each covering a 32×32-pixel region.
Key methods: spatially joint reconstruction via Pixel Shuffle + 3×3 convs (fixing U1's patch-boundary artifacts from independent token decoding), resolution-aware noise conditioning extended to 4096×4096, and a 'specialize, then unify' scheme — four RL experts (aesthetics, bilingual text rendering, infographics, editing) consolidated via multi-expert on-policy distillation. Training added 59M text-image pairs from 78 sources and 38M editing examples, with 88.2% of effective generation volume above 1024². Known issues remain with dense small text, small faces/hands and multi-reference drift. Open-sourced on GitHub and Hugging Face.
More from Multimodal
- Musubi Tuner merges MiniMax H3 video-audio model training support — bdsqlsz · 2026-09-16
- Unfold wins two awards at OpenAI Singapore hackathon for turning 2D manuals into 3D guides — cedric_chee · 2026-09-16
- Unverified GPT-6 Astra demo rebuilds indoor scenes in 3D from one reference image — Dazzling_Gate_330 · 2026-09-16
- Creator recreates viral volcano popcorn meme entirely with flovaai — SimplyAnnisa · 2026-09-16
- Creator generates 30-second AAA gameplay concept with Seedance 2.5 on Runway — azed_ai · 2026-09-16
- Creator makes 30-second AAA game concept with Seedance 2.5 on Runway — azed_ai · 2026-09-16