Persistent representation learning cuts Drifting Models FID by 82–95% from pixels
Doudou Zhang · hf · 2026-10-07
- Context: Drifting Models shift iterative distribution refinement from inference to training for one-step generation, but performance hinges on the representation used to build the drifting field — pixel-space fails while pretrained feature spaces excel, for unclear reasons.
- Finding: the gap traces to the discriminative geometry of the representation, which governs sample weighting in KDE and thus drift; the authors prove a current-step gradient equivalence between the KDE ratio loss and drift regression loss, motivating direct control of drifting velocity.
- Method: persistent representation learning continuously learns a more discriminative representation as the generator evolves across batches.
- Results: across datasets, learning from pixels alone reduces FID by 82–95% over original pixel-space Drifting Models, without pretrained encoders; adapting pretrained representations and velocity clipping add further gains.
More from Multimodal
- VEDA Sparse Attention cuts MiniMax H3 video gen time in half in ComfyUI with no visible quality loss — robomar_ai_art · 2026-10-07
- a16z consumer AI overview turned into a 3-minute video with Manus — parker_lyman · 2026-10-07
- Blogger says Opus-generated video matches museum-grade immersive work that once cost $70K teams — oran_ge · 2026-10-07
- Super robot anime made with Kling 4 Flash looks straight out of a studio — Forsaken_Stuff_Ai · 2026-10-07
- Dev builds ComfyUI node pack for projection mapping, 3D region prompts and motion design — Puzzled_Parking2556 · 2026-10-07
- Full Nano Banana 2.1 Prompt for Cinematic Posters That Keep Your Face and Copy the Layout — CodeByPoonam · 2026-10-07