SenseTime Launches Native 8K Image Generation Model

量子位 · wechat · 2026-07-18

At WAIC 2026, SenseTime launched a new model, **日日新 SenseNova U1 Pro**. The article describes it as a "delivery system" for complex multimodal tasks—capable of understanding, planning, organizing information, generating multimodally, and self-correcting to deliver a final product, rather than just generating images. Key demonstrated capabilities include **native 8K ultra-long image output**, continuous interleaved text-and-image creation, and the ability to deliver finished assets for infographics, urban planning, film storyboards, academic posters, and commercial design. The author tested it with a 24 solar terms scroll, a lab-style poster, a glazed landscape painting, and a movie poster, noting solid performance in long-image details, layout stability, and complex style execution. Technically, SenseTime reportedly uses **32×32 large patches** to manage the visual token and VRAM pressure of 8K, combined with adaptive NoiseControl, patch overlapping, spatial sampling, and specific loss designs to preserve small text, textures, and structures. The article concludes that multimodal generation is shifting from "single-point image output" to "system-level content delivery," with U1 Pro representing an early form of this trend.

Related event: SenseTime Launches SenseNova U1 Pro Native 8K Multimodal Model(6 posts)→

Original post →

More from Companies & People

Companies & People channel →