WaiT for the Signal adds frequency-aware flow matching and cuts sampling compute by 50%
TimDarcet · x · 2026-08-04
WaiT for the Signal proposes a frequency-aware flow-matching approach for image generation.
- It incorporates image frequency structure directly into diffusion/flow models by decomposing generation into coarse and fine bands with lossless wavelets.
- High-frequency bands are kept as noise until coarse structure emerges, then refined jointly.
- The authors also introduce a stricter three-axis evaluation protocol at native resolution, arguing that standard FID hides fine detail.
- Results include pixel-space FID 1.43 on ImageNet 512×512, up to 50% less sampling compute, and a new 1.3 FID for a 2B pixel-space model; the method also reports SOTA FVD 0.84 on Kinetics-600 without algorithmic changes.
Related event: WaiT Model Uses Frequency-Aware Flow Matching to Cut Sampling Compute(2 posts)→
More from Multimodal
- Fizgig adds experimental LoRA training for MiniMax H3 with 15.7 GB model files — shootthesound · 2026-08-04
- MiniMax H3 adds open weights, stereo audio and 15-second 2K video generation — petrusenko_max · 2026-08-04
- MiniMax H3 turns an Office prompt about coding agents into a one-shot demo — venturetwins · 2026-08-04
- Using Claude with prebuilt shaders and assets to generate game worlds — mathemagic1an · 2026-08-04
- invideo’s Agent Two adds notebooks for finer control over image, video and audio — azed_ai · 2026-08-04
- MiniMax H3 day-one test finds better identity retention with max ref image size — jozbgm · 2026-08-04