Salt++ post-training lifts 4-step streaming audio-video generation by up to 57%
Xingtong Ge · hf · 2026-10-08
Salt++ is a two-stage post-training framework for few-step streaming audio-video generation. Causal Self-Flow aligns a noise-mixed-history student with a clean-history EMA teacher to improve semantic extraction and cross-modal alignment, while context-aligned autoregressive DMD shares causal masks and prefixes across sampling and score training. On JavisBench at 480p it improves visual and motion quality by 57% and 45% over OmniForcing under the same 4-step causal setting, and a scale-wise stage extends it to 1664×960, beating bidirectional LTX-2 on six of seven metrics.
More from Multimodal
- Nano Banana 2.1 Image Model Available for Free via Google Flow — aziz4ai · 2026-10-08
- Nano Banana 2.1 image generation available for free in Google Flow — aziz4ai · 2026-10-08
- Local LLM + Tripo3D: A Mini 4WD-Style 3D Racing Game Built in One Day — lxfater · 2026-10-08
- ChatGPT finds 10 anatomical failures in a GPT Image 2.5 generation — cocosus · 2026-10-08
- Local ComfyUI workflow: Qwen 2.1 nails product references but fails on people — Busy_Site_8767 · 2026-10-08
- Google's Flow Music Plugin Generator Turns Text Prompts Into Audio Effects Plugins — teropa · 2026-10-08