SenseNova u1.5 Architecture Analysis: Can ConvDecoder Eliminate Grid Artifacts?

Taylar214 · reddit · 2026-08-11

The author provides an in-depth analysis of the SenseNova u1.5 architecture update. The new version introduces a ConvDecoder that converts visual tokens into a 2D grid and upsamples them using pixel shuffle and 3x3 convolutions. This allows neighboring patches to communicate during image reconstruction, aiming to fix grid artifacts and broken textures at super high resolutions.

However, the author points out that the official image generation scores conflate multiple variables (new decoder, more training data, better prompts, etc.) without a proper ablation study. The author calls for more rigorous testing:

Furthermore, the author notes that despite complex structured formats being rare in the generation/editing training data, the model still excels at following long, structured prompts. This suggests a potential cross-task transfer where structural understanding learned from comprehension tasks aids visual planning, which could be a significant finding if proven via ablation.

Original post →

More from Multimodal

Multimodal channel →