Sony AI open-sources PAVAS, a physics-aware video-to-audio model accepted as CVPR 2026 Oral

mittu1204 · x · 2026-10-03

PAVAS, a physics-aware video-to-audio synthesis model led by Sony AI intern Hyun-Bin Oh under Yuhta Takida's guidance, has been accepted as a CVPR 2026 Oral, and the source code is now public on GitHub.

Built on MMAudio (CVPR 2025), PAVAS augments the generation backbone with object-centric conditioning derived from mass, velocity, segmentation, and patch-level visual features, so generated audio better reflects the physical interactions in a video. The repo includes the model and training code (pavas/core), a training-free Physics Parameter Estimator (PPE) with offline precompute stages, and training/evaluation entrypoints.

The paper PDF and online demo are available, and the team invites the community to include PAVAS in CVPR 2027 benchmarks.

Original post →

More from Multimodal

Multimodal channel →