Qwen presents OmniVChat: synthesized data, benchmark and RL for native audio-visual dialogue

Qwen · hf · 2026-09-21

Qwen's team introduces OmniVChat, the task of native audio-visual dialogue where an omni model directly consumes simultaneous audio+video input (queries embedded in the stream, no separate text, captioning, or ASR), cutting latency while preserving perceptual cues. To tackle data scarcity and evaluation, the team builds OmniVChat-Studio, a multi-agent engine for synthesizing single- and multi-turn dialogues, and OmniVChat-Bench spanning five ability categories (plus a human-recorded variant). They also propose OmniVChat-RL, a reward design jointly targeting correctness, efficiency, and style. Training Qwen3-Omni-Instruct on synthesized dialogues improves scores on both benchmarks, validating the reward design and showing transfer to real-world conversations.

Original post →

More from Multimodal

Multimodal channel →