NVIDIA Releases Open-Source AV-Flamingo Audio-Visual LLM
nvidia · hf · 2026-07-20
NVIDIA has released the fully open-source audio-visual large language model Audio-Visual Flamingo (AV-Flamingo), specializing in understanding and reasoning over long, complex real-world videos and audio.
Key highlights:
- Training Data & Methods: Built a dataset of roughly 7 million caption and QA instances, employing a three-stage curriculum learning approach that transitions from short-term perception to long-horizon, multi-event reasoning.
- Reasoning Framework: Introduces a temporal audio-video interleaved chain of thought, explicitly anchoring reasoning steps to timestamps in the audio-visual stream to improve temporal alignment and interpretability.
- Performance: Across 15+ audio-visual and omni-modal benchmarks, it not only significantly outperforms open-source models of similar scale but even surpasses larger closed-source models on certain long-video tasks.
More from Multimodal
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22