NVIDIA Releases Open-Source AV-Flamingo Audio-Visual LLM
nvidia · hf · 2026-07-20
NVIDIA has released the fully open-source audio-visual large language model Audio-Visual Flamingo (AV-Flamingo), specializing in understanding and reasoning over long, complex real-world videos and audio.
Key highlights:
- Training Data & Methods: Built a dataset of roughly 7 million caption and QA instances, employing a three-stage curriculum learning approach that transitions from short-term perception to long-horizon, multi-event reasoning.
- Reasoning Framework: Introduces a temporal audio-video interleaved chain of thought, explicitly anchoring reasoning steps to timestamps in the audio-visual stream to improve temporal alignment and interpretability.
- Performance: Across 15+ audio-visual and omni-modal benchmarks, it not only significantly outperforms open-source models of similar scale but even surpasses larger closed-source models on certain long-video tasks.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11