NVIDIA Releases Open-Source AV-Flamingo Audio-Visual LLM
nvidia · hf · 2026-07-20
NVIDIA has released the fully open-source audio-visual large language model Audio-Visual Flamingo (AV-Flamingo), specializing in understanding and reasoning over long, complex real-world videos and audio.
Key highlights:
- Training Data & Methods: Built a dataset of roughly 7 million caption and QA instances, employing a three-stage curriculum learning approach that transitions from short-term perception to long-horizon, multi-event reasoning.
- Reasoning Framework: Introduces a temporal audio-video interleaved chain of thought, explicitly anchoring reasoning steps to timestamps in the audio-visual stream to improve temporal alignment and interpretability.
- Performance: Across 15+ audio-visual and omni-modal benchmarks, it not only significantly outperforms open-source models of similar scale but even surpasses larger closed-source models on certain long-video tasks.
More from Multimodal
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11