Hi-DiT hybrid Latent-Pixel diffusion hits FID 1.06 on ImageNet, accepted at ECCV 2026
量子位 · wechat · 2026-09-26
Researchers from USTC and HiDream.ai propose Hi-DiT (Hybrid Latent-Pixel Diffusion Transformer), accepted at ECCV 2026.
- A shared Transformer backbone coordinates a Latent stream (global structure in high-noise phases) and a Pixel stream (high-frequency detail in low-noise phases), with time-gated injection at τ=0.3.
- A High-Frequency Pixel Predictor uses conv + PixelShuffle to hierarchically decode sub-patches, easing high-dimensional pixel regression.
- Results: FID 1.06 on ImageNet 256×256 (CFG), gFID 1.26 on 512×512; ablations show pure Pixel stream gets 3.38, +Latent input 1.74, +predictor head 1.65.
- Overhead is minimal: 17.7→18.1 min/epoch, 1.58→1.67 s/image, 35.52GB peak memory. Code open-sourced at github.com/HiDream-ai/Hi-DiT.
More from Multimodal
- PrunaAI's distilled Qwen-Image-2.1 with few-step generation trends on Hugging Face — PrunaAI · 2026-09-27
- Dev's tested AI music workflow: lyrics, Suno, ear-curation, then Ableton — ctjlewis · 2026-09-27
- Tailored ASR for Japanese speaking assessment cuts mora error rate from 12.3% to 7.1% — tkasasagi · 2026-09-27
- A full AI music video now costs ~$65 and 6M tokens — and it's no longer special — rickasaurus · 2026-09-27
- One Prompt, a 60-Second Singularity Video Essay: Runway CEO Demos Agentic Video Editing — c_valenzuelab · 2026-09-27
- Redditor argues Krea 2 is still the best full HD image model, ahead of its time — Due_Research9042 · 2026-09-27