Salt++ post-training lifts 4-step streaming audio-video generation by up to 57%

Xingtong Ge · hf · 2026-10-08

Salt++ is a two-stage post-training framework for few-step streaming audio-video generation. Causal Self-Flow aligns a noise-mixed-history student with a clean-history EMA teacher to improve semantic extraction and cross-modal alignment, while context-aligned autoregressive DMD shares causal masks and prefixes across sampling and score training. On JavisBench at 480p it improves visual and motion quality by 57% and 45% over OmniForcing under the same 4-step causal setting, and a scale-wise stage extends it to 1664×960, beating bidirectional LTX-2 on six of seven metrics.

Original post →

More from Multimodal

Multimodal channel →