单卡 5090 从零训练图像模型:冻结编码器突破 VAE 模糊瓶颈
Creative-Listen-6847 · reddit · 2026-07-24
A developer is sharing their ongoing journey of training a DiT-based image generation model from scratch on a single RTX 5090. Constrained by the 32GB VRAM, training is locked at 512x512 resolution. Pushing the resolution to 1024x1024 causes severe structural breakdown—not because of the DiT, but due to the underlying VAE hitting its decoding ceiling, amplifying latent blurs and noise.
To fix this without wasting hundreds of hours of generator training, the author devised a clever workaround: freeze the VAE encoder to keep the latent space intact, and only train the decoder to output sharper images.
「多模态」频道最新
- Seedance 2.0 出现在 CapCut,展示一段高完成度动效作品 — bennash · 2026-07-24
- Midjourney 的 `--draft --sref random` 一次可出 24 种风格方向 — ciguleva · 2026-07-24
- 辟谣:疯传的 FLUX 3 预告系 AI 生成,非官方发布 — gandamu_ml · 2026-07-24
- 一个 ComfyUI 节点能把生成图自动上传到可分享 moodboard — massivebacon · 2026-07-24
- 一个简单的 ComfyUI `supir_restore` 工作流把 420×460 放大到 2159×2920 — Visible_Motor_3138 · 2026-07-24
- 一个 Midjourney 提示词把任意主体变成蓝图示意图 — tisch_eins · 2026-07-24