H3 Audio Generation Workflow: Achieving 48kHz High-Fidelity Upsampling via AudioSR

AppointmentOk972 · reddit · 2026-08-11

The author shares an audio generation and optimization workflow based on the H3 model. By modifying the official T2VA workflow, it uses a 32x32 image dimension input to minimize video generation time and bypasses the video VAE.

The core trick is routing the Audio VAE output into an AudioSR custom node to achieve high-quality 48kHz upsampling, ultimately saving as FLAC. The author notes that the dynamics and sound field are highly impressive and recommends integrating this AudioSR upsampling trick into standard video workflows.

Original post →

More from Multimodal

Multimodal channel →