EG-FM Lowers FID to 1.45 in Pixel-Space Image Generation Without Backbone Changes
burny_tech · x · 2026-08-10
The paper Energy-Guided Flow Matching (EG-FM) introduces a new training path for pixel-space generative models.
- Core Improvement: Instead of standard flow matching that pushes noise toward a fixed clean-image endpoint, EG-FM uses a moving endpoint. Based on the image's spectral energy, it adopts a coarse-to-fine trajectory—starting with low-frequency global structures before gradually releasing high-frequency details.
- Zero-cost Integration: The framework requires no changes to the backbone architecture or training data, bringing negligible costs to both training and inference stages.
- Performance: It achieves consistently lower FID on ImageNet 256x256 with fewer epochs (1.55 at 200 epochs, 1.45 at 600 epochs). For 512x512 high-resolution adaptation, it yields an FID of 1.58 after only 40 steps. Transferred to text-to-image generation, it scores 0.85 on GenEval and 83.9 on DPG-Bench.
Related event: EG-FM Method Enhances Pixel-Level Image Generation Efficiency(2 posts)→
More from Multimodal
- Decart AI launches Anywear Chrome plugin for real-time virtual try-on — sophiamyang · 2026-08-10
- Character LoRA Fails in Krea2 Raw Mode, Developer Seeks Workflow Help — maxiedaniels · 2026-08-10
- Minimax H3 Impresses in Video Accuracy, Devs Explore Reference-to-Image Workflow — ColdExample · 2026-08-10
- Runway Video Generation Test: Cinematic Lighting and Textures — Eric520CC · 2026-08-10
- Creator Combines Seedance 2.5 and Other AI Models for Emotional Short Film — darkshark9 · 2026-08-10
- AI Video Generation Tested: Epic Fantasy Battle Scenes Yield Impressive Results — umesh_ai · 2026-08-10