Generating Low-Res Flat Color Images
PhatTipAndRawNips · reddit · 2026-07-17
The author aims to tackle a highly specific text-to-image task: generating images at a mere 10×10 to 30×30 pixels resolution while maintaining flat colors, high contrast, and a recognizable subject, alongside fast generation speeds and low costs.
They experimented with several approaches:
- Retro Diffusion: Capable of generating tiny images, but naturally leans towards gradients and excessive colors. Even with palette restrictions, the results are often suboptimal.
- SD-piXL: Works well on larger images and supports color limits, but quality degrades severely at minuscule sizes, and it runs slowly even on an RTX 4090.
- Downscaling high-res images / applying k-means color reduction: Results in the loss of crucial details needed to keep the subject recognizable.
The author is seeking better community solutions and will only consider training a custom model as a last resort.
More from Multimodal
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- Same prompt, Seedance 2 and Grok are compared on cinematic transformation output — LudovicCreator · 2026-07-21
- CG Chefs Showcases Retro Anime Style AI Video Generation — nicolascraske · 2026-07-21
- Night-party video demo uses Seedance 2.0, timecode prompts and 4K upscaling — gen_ericai · 2026-07-21