Data-space loops beat uniform flow schedules at ImageNet FID 2.46

Few-Step Generation via Data-Space Iteration

Shanchuan Lin, Yansong Peng, Fu-Yun Wang, Haoqi Fan

cs.LG, cs.CV

2026-10-08

Under one DMD recipe, data-space iteration refines a single prediction for four steps and records ImageNet 256 FID 2.46, ahead of uniform-grid flow-renoise at 2.75.

What problem this solves

Flow matching learns a velocity field from noise to data. Sampling integrates that field, often for tens or hundreds of network evaluations. Distillation can collapse the integral to one pass, but one forward has too little depth and the images get visibly worse. The operating point that matters is a few steps. Deployed high-resolution systems often use four.

Existing few-step generators still iterate on the probability flow. Flow-renoise predicts a clean image, then mixes noise back in at the next scheduled time. Flow-map methods, including phased DMD, learn the jump between two scheduled times. Both require a timestep grid fixed before training. With self-rollout, each update trains on states the generator itself just produced, so a new grid means a new distilled model. The grid is global. Class and spatial frequency change how hard refinement is, and the step sizes stay locked to one table.

Method

Data-space iteration leaves the time axis. Noise z is drawn once and held fixed. The prediction starts at 0. A shared network repeatedly emits z minus a velocity. Inputs are the previous data prediction, the same z, and a step index r, with flow time pinned at 1. Steps do not re-noise, and they do not build intermediate flow states. Every step is trained to return the best sample its capacity allows. Conditioning on r lets the correction change by stage. Dropping r raises four-step FID from 2.46 to 2.63.

Training swaps only the DMD rollout. A frozen teacher velocity and a fake velocity that tracks the student, also conditioned on r, supply the update direction. The first r steps run with stopped gradients; only the current step backpropagates. Compute and memory match a DMD baseline on the same recipe. GAN and rectified-flow auxiliaries stay out, so iteration space is the variable under test. Prior self-conditioning still integrates an ODE and uses the previous prediction only to correct velocity. Here the iterated state is the data prediction.

The generator starts from the pretrained velocity network. An extra projection of the current prediction is added to the noise projection, and an embedding of r joins the time pathway. Flow time is fixed at 1. The output is z minus the predicted velocity, the usual data read-out at the noise end of a linear path. Replacing the timestep MLP with a constant embedding cuts the count from 675,129,632 parameters to 673,527,200.

The loop looks like a fixed-point iteration. It does not have to converge. The index r allows a different map at each step, and the distribution-matching loss hits every prefix, not only a converged state. An optimal-transport penalty at λ of 1e-6, meant to pull every step toward one transport, moves four-step FID from 2.46 to 2.47 and is dropped.

Results

The test is class-conditional ImageNet 256 in latent space, with an SiT-XL/2 teacher. FID compares 50k class-balanced samples to the training set. Every run uses R of 4, learning rate 1e-4, batch size 256, and 40 epochs, with four fake-model updates per generator update. Adam at β1 of 0 and β2 of 0.95 reaches four-step FID 2.46; β1 of 0.9 reaches 2.66. DMD draws τ from Beta(2.5, 1) with weight 1. Uniform τ at the same weight lands at 3.61. Two classic score weights that diverge as τ approaches 0 simply crash.

Four-step FID without guidance:

MethodTime gridFID
SiT teacher, 250 NFEODE8.26
DMD flow-renoiseuniform2.75
DMD flow-renoisebest of 3 shifts2.72
Phased DMDuniform2.55
Phased DMDbest of 3 shifts2.48
Data-space iterationnone2.46

Data-space FID at 1, 2, 3, and 4 steps is 3.85, 2.68, 2.47, and 2.46, flat after step 3. Flow-renoise goes from 2.67 at 3 steps back to 2.75 at 4. On the uniform grid, data-space is lower at every NFE. Flow-map intermediate states are not valid samples, so the table has only the fourth step.

CFG at scale 1.1, restricted to τ in [0, 0.6], is the best guidance setting: data-space 2.29, flow-renoise 2.61, uniform phased DMD 2.86, and the best of three phased DMD schedules 2.37. Scale 1.1 over the full τ range lifts four-step data-space FID from 2.46 to 2.59; scale 1.2 lifts it to 3.27. The guided teacher needs 250×2 evaluations and reaches 2.06. Shift factors 0.5, 1, and 2 swap rank across methods and guidance, so the winning grid is not predictable.

Four unguided steps do not win every metric. Data-space sFID is 4.48 and recall is 0.59, better than flow-renoise at 4.62 and 0.55 and phased DMD at 5.05 and 0.56. Inception Score goes the other way, 245.76 versus 228.59, and precision is 0.83 to 0.84 versus 0.81. Mean RGB L2 after decoding, pixels scaled to [0, 1] over 50k images, is 82.72 versus 59.65 on the first interval. On the third interval data-space still moves 35.98 while flow-renoise moves 29.91. Mixing noise back in wipes regions that were already coherent.

On CIFAR-10, a 55M SongUNet with no extra tuning, the teacher FID is 3.54 at 250 steps. At four steps, data-space reaches 8.62, flow-renoise 10.03, and phased DMD 11.72. At one step the rollouts reduce to the same formula, and data-space is slightly worse, 12.20 versus 11.21. Three runs of data-space and flow-map have FID standard deviation below 0.05.

Why it matters

In a DMD pipeline the change is the rollout. One noise sample rewrites a data prediction, and there is no extra student per timestep grid. Without guidance the uniform-grid baselines are 2.75 and 2.55 against 2.46. The best of three hand-picked grids is 2.48, an edge of 0.02. With guidance, unsearched phased DMD is 2.86, the searched run is 2.37, and data-space is 2.29.

The quality gain is incremental. What gets removed is a training search where each grid is another model. GAN and rectified-flow terms were left out on purpose, so 2.46 is not a public-leaderboard number. The saving is useful where self-rollout is already required, as in causal streaming video. If the grid is chosen only at inference, the case is weaker.

Limitations

Self-rollout is required. Distillation that already rolls out pays nothing extra. For a pointwise ODE objective, paired data and noise can train the loop, and a cheaper recipe is still open. The main result is ImageNet 256. Text-to-image is future work. The paper states that more than four steps does not lower FID further.

Beating a searched schedule is not stable without guidance. The gap from 2.46 to 2.48 is 0.02, and the retraining standard deviation is below 0.05. The guided gap, 2.29 versus 2.37, clears that band. The search is three shift values. If DMD times are drawn only from the current phase up to 1, one-step FID jumps to about 44, while four-step FID stays near 2.5. The sampling grid is gone. The loss still depends on the time distribution.

The introduction says different samples and spatial locations want different step sizes. That was not measured. The index r is global. Removing the grid is not evidence of spatially adaptive refinement. Unguided DMD students beat the 250-step teacher by a wide margin. The paper guesses this comes from mode-seeking of the reverse KL used by DMD. Recall falls from 0.67 to 0.59. With guidance, four-step FID 2.29 still trails the teacher at 2.06.

Terms

Source

What people are saying

Related papers

All paper explainers