Qwen presents OmniVChat: synthesized data, benchmark and RL for native audio-visual dialogue
Qwen · hf · 2026-09-21
Qwen's team introduces OmniVChat, the task of native audio-visual dialogue where an omni model directly consumes simultaneous audio+video input (queries embedded in the stream, no separate text, captioning, or ASR), cutting latency while preserving perceptual cues. To tackle data scarcity and evaluation, the team builds OmniVChat-Studio, a multi-agent engine for synthesizing single- and multi-turn dialogues, and OmniVChat-Bench spanning five ability categories (plus a human-recorded variant). They also propose OmniVChat-RL, a reward design jointly targeting correctness, efficiency, and style. Training Qwen3-Omni-Instruct on synthesized dialogues improves scores on both benchmarks, validating the reward design and showing transfer to real-world conversations.
More from Multimodal
- Runway Big Pitch Contest entry 'ALL THE WATER' imagines oceans vanishing underground — bennash · 2026-09-21
- MiniMax H3 Video: How to Use a Second Image to Replace a Body Part — Next-Place0 · 2026-09-21
- Paper: Physically Based Rendering in the Latent Space — ssh4net · 2026-09-21
- Adaptive Color Grading paper: KNN beats end-to-end models at tonescale prediction — ssh4net · 2026-09-21
- Blender as Director, Seedance as Renderer: A Cinematic AI Video Workflow — CurieuxExplorer · 2026-09-21
- MiniMax H3 Video Model Passes Fast-Cut Montage Test With Consistent Characters — Hailuo_AI · 2026-09-21