Sony AI open-sources PAVAS, a physics-aware video-to-audio model accepted as CVPR 2026 Oral
mittu1204 · x · 2026-10-03
PAVAS, a physics-aware video-to-audio synthesis model led by Sony AI intern Hyun-Bin Oh under Yuhta Takida's guidance, has been accepted as a CVPR 2026 Oral, and the source code is now public on GitHub.
Built on MMAudio (CVPR 2025), PAVAS augments the generation backbone with object-centric conditioning derived from mass, velocity, segmentation, and patch-level visual features, so generated audio better reflects the physical interactions in a video. The repo includes the model and training code (pavas/core), a training-free Physics Parameter Estimator (PPE) with offline precompute stages, and training/evaluation entrypoints.
The paper PDF and online demo are available, and the team invites the community to include PAVAS in CVPR 2027 benchmarks.
More from Multimodal
- GemPix 2.5 Flash string spotted in Gemini iOS update, hinting at Nano Banana 2.5 Flash — lyraxana · 2026-10-03
- PotionUI 0.0.14 Drops the GPU Requirement: Cloud Models via OpenRouter, Built-in Editor, Video Director — 0roborus_ · 2026-10-03
- Researcher trains a diffusion model from scratch in two weeks — and finds its art beautiful — pbaylies · 2026-10-03
- Static: a sci-fi fake trailer made entirely with Kling — SightsFilms · 2026-10-03
- Redditor Releases SITCOM, an AI-Generated Psychological Horror Short Film — Sufficient_Flow_415 · 2026-10-03
- Glitched Gotham: a Midjourney --sref code for glitch aesthetics — egeberkina · 2026-10-03