Velocity Scaling in Flow Matching
Youssef Saied, François Fleuret
cs.CV
2026-10-08
Multiplying a flow velocity by a lag-calibrated gain, without retraining, cuts unguided SiT-XL/2 ImageNet-256 FID from 28.0 to 12.2 at NFE 25, and a sweep reaches 8.7.
Flow matching pushes noise to data along a straight path by integrating a learned velocity. With few steps the integrator skips the updates that grow sample-specific structure, and the state falls behind the clock the solver is using.
Li et al. (ICML 2026) multiply the predicted velocity by a gain γ(t). On unguided SiT-XL/2 at 50 network evaluations, FID drops from 13.7 to 7.6. They trace the drop to mean-squared-error training shrinking velocity norms, and propose a Scale Schedule Corrector, γ(t) = 1.1 − 0.1t. Schedules with a similar average gain land between FID 7.5 and 7.8.
The population minimizer is the conditional mean v(x, t) = E[U | Xt = x], and that field transports the source onto the data law exactly. The mean is shorter than a single target because several training paths cross at one state. The residual has conditional mean zero, so it does not move the distribution. On a Gaussian path, gain 1 reaches the target covariance Σ; any other constant gain yields Σ^γ. A pretrained SiT, scored at 20 times on 16,000 ImageNet image-noise pairs, has a best-fit scalar multiplier whose mean is 0.999. On genuine training paths there is no missing velocity magnitude to restore.
What scaling actually closes is population time lag: at solver time t, generated states read as earlier points on the training path. Figure 1 holds the class, the initial noise, and the Brownian path fixed. At 25 evaluations the unscaled object forms later than at 100; constant gain 1.15 pulls the coarse trajectory forward. A two-step Euler state at t = 0.5 lacks per-image structure. A scale-normalized probe reads it as time 0.254, against 0.493 for a genuine state.
The probe trains only on genuine paths. A three-layer conv net emits the posterior mean over 20 time bins. Held-out mean absolute error is 0.0258 for SiT, 0.0250 for the ADM U-Net, and 0.0141 for CIFAR-10. Lag is the gap between mean readings on genuine and sampled states, so probe bias is not counted as lag. Single states are noisy: the SiT probe sits 0.0226 mean absolute deviation off its population curve. Moving a genuine state 0.05 back along its path drops the estimate by 0.0499; an equal perpendicular move drops it by 0.00519.
Model-time correction leaves velocity untouched and feeds the network t minus the measured lag. Velocity scaling leaves the time input untouched and multiplies velocity by γ, so the state catches the clock before the next call. To first order, a step of length h that should cover a lag τ − t wants γ(τ − t) ≈ h.
The lag-based gain γlag is the constant that minimizes mean squared lag along trajectories, at resolution 0.01, with 256 trajectories per candidate and no image metric. For SiT-XL/2 at 50 evaluations the search is 6 candidates and 76,800 forward passes. One 50,000-image FID point at that budget is 2.5 million forwards, a median 4.7 RTX 4090 hours. At matched average gain 1.07216, a fitted schedule scores FID 13.0 and a constant scores 12.5.
Unguided ImageNet-256, 50,000 paired images per cell, uniform stochastic Euler. Lag-gain FID is linearly interpolated. The abstract writes 12.21; the table rounds to 12.2. No uncertainty interval is reported for those summaries. Li et al.'s 13.7 comes from a different hybrid SDE and should not be lined up against the 28.0 below.
| Model | NFE | Unscaled | Lag gain | Best on grid |
| SiT-S/2 | 25 | 71.1 | 50.9 | 44.4 (γ=1.15) |
| SiT-B/2 | 25 | 47.9 | 30.2 | 25.9 (γ=1.125) |
| SiT-L/2 | 25 | 30.2 | 16.1 | 13.4 (γ=1.125) |
| SiT-XL/2 | 25 | 28.0 | 12.2 | 8.7 (γ=1.15) |
| ADM U-Net | 25 | 34.8 | 17.9 | 14.7 (γ=1.125) |
| SiT-XL/2 | 250 | 9.7 | 9.1 | 6.3 (γ=1.05) |
| ADM U-Net | 250 | 16.8 | 15.8 | 12.8 (γ=1.05) |
Across twenty ImageNet settings the lag gain recovers 49% of the FID drop from gain 1 to the best grid gain: 81%, 63%, 38%, and 15% at 25, 50, 100, and 250 evaluations. On CIFAR-10 the decoder-free NCSN++ prefers gain 1. At 50 evaluations the interpolated lag gain is worse, 4.1 against 3.0 unscaled.
Changing only the time input, on the public hybrid SDE at 50 evaluations, cuts FID from 13.8 to 8.8, or to 8.6 after global-norm normalization. The same correction on an ADM U-Net goes from 22.7 to 16.7. Constant gain 1.075 reaches 13.2.
A frozen-field control replays velocities recorded on the unscaled path and never re-evaluates the moved states. Across five SiT checkpoints it keeps only 22.4% to 24.9% of the full FID drop under stochastic sampling. Under deterministic sampling the share rises to 53.9% for SiT and 59.0% for the ADM U-Net. A radial extra push, with the network still re-evaluated, makes FID worse.
The best grid gain sits above γlag in every one of the twenty settings. A multiplier κ fit at 25 evaluations transfers to higher budgets with mean absolute gain error 0.011 across four SiT sizes. At 250 evaluations, κ gives SiT-XL/2 FID 6.6 against a grid best of 6.3. At a matched 50 evaluations, stochastic Heun lowers lag RMS from Euler's 0.019 to 0.005 and the best gain from 1.10 to 1.05. Stronger classifier-free guidance pushes the best gain below 1. The accompanying rise in velocity norm is too small to account for that shift. FID prefers a higher gain than sFID, and that higher gain raises precision while lowering recall. Scaling also assigns a larger share of endpoint variance to directions that already vary more in real data. Gains matched on that sharpening land within 0.2 FID of the best measured value at 50, 100, and 250 evaluations.
The gain is an inference-time multiplier. Weights stay frozen. At 25 evaluations, SiT-XL/2 moves from FID 28.0 to 12.2 with the lag gain and to 8.7 at gain 1.15. Calibration costs about 77,000 forwards. One 50,000-image FID curve costs a median 4.7 GPU-hours. The lag gain recovers about four fifths of the available drop at 25 evaluations and about 15% at 250. Past that, the recipe is κ or a metric sweep.
A fixed factor near 1.1 is the wrong default once the step count rises, the solver becomes Heun, or guidance gets strong. The preferred gain moves toward 1, and under strong guidance it falls below 1. The CIFAR-10 pixel model wants gain 1.
The failure mode is a state that does not match its time. MixFlow trains on earlier states, the Time-Shift Sampler edits the next time input, and Epsilon Scaling shrinks the noise prediction inside diffusion. Velocity scaling advances a lagging state along the learned velocity, then asks the network for a new velocity there. A magnitude deficit fits neither an exact population field nor a measured multiplier of 0.999.
Lag does not account for the whole FID drop. In all twenty unguided stochastic ImageNet settings the best grid gain is higher, including at 250 evaluations where the measured lag is already small. κ, fit at 25 evaluations, extrapolates with mean absolute error 0.011. That is a stable ratio, not two separate mechanisms, and it supplies no limiting gain for still finer integration.
The headline 12.2 is interpolated. Two independent 50,000-image runs at gain 1.07216 score 12.5 and 12.7, so run-to-run scatter is about 0.2. That scatter is far smaller than the gap from 12.2 to 8.7, and the tenths between neighboring grid points have no reported interval. The probe cannot correct a single image. Three training seeds agree on gain 1.04 for SiT-XL/2 at 50 evaluations and resolution 0.01, but the bootstrap intervals freeze the probe.
Matching endpoint moments after sampling does not reproduce the gain. At 50 evaluations, rescaling radius leaves FID at 13.8. At 100 evaluations, matching the full latent mean and covariance of a gain-1.05 run moves FID from 11.6 to 12.3. At 250 ODE evaluations, unscaled endpoint mean squared error is 0.004 and SSC stays at 0.052. SSC also increases one-step error. How much of the drop needs a fresh velocity depends on the sampler: the frozen field explains about a quarter under stochastic sampling and about half under deterministic sampling. Coarse discretization does create lag. It does not exhaust the effect.
On CIFAR-10, gains above 1 do not help, and a lag gain slightly above 1 hurts. A REPA checkpoint is inside the experimental scope, but the main tables do not list it on its own. Code is described as forthcoming. Under strong guidance, the measured increase in velocity norm does not cover how far the best gain falls below 1. A magnitude-aware loss inflates the conditional mean and leaves the residual where it was. Image metrics for that loss are not reported.