CUHK and Alibaba propose OmniVChat: native audio-visual dialogue trained on synthetic data
机器之心 · wechat · 2026-09-28
Researchers from CUHK, Alibaba TokenHub, SJTU and others introduce OmniVChat, defining native audio-visual dialogue where a model directly consumes user audio and video simultaneously — no text prompts, subtitles, or ASR needed — preserving tone, expressions, and camera orientation.
To tackle scarce real dialogue data and missing evaluation standards, they take a 'generation for comprehension' route with three components: a multi-agent data engine (Studio) synthesizing single/multi-turn dialogues with reference replies and tiered rubrics; a 2,800-case benchmark (Bench) covering connection awareness, entity alignment, self-awareness, anti-hallucination, and emotion recognition; and an RL scheme (OmniChat-RL) turning rubrics directly into rewards.
Built on Qwen3-Omni-30B-A3B-Instruct, the key finding: RL on purely synthetic data lifted rubric scores from 0.465 to 0.652 on Bench, and from 0.402 to 0.632 on a human-recorded set never seen in training — synthetic training generalizes to real interactions. Paper, dataset, and code are open-sourced.
More from Research
- ZJU Releases EMem-Bench: 2,554 Episodes to Test Embodied Agent Memory — OmniAI-ZJU · 2026-09-29
- Tencent's AdaTutoRank Trains Setwise RAG Rerankers via Adaptive Tutoring Optimization — tencent · 2026-09-29
- Calibrated Importance Sampling Fixes Training-Inference Mismatch in LLM RLVR — Tianrun Yu · 2026-09-29
- Stanford lands $25M NIH grant to build AI center for personalized dementia care — StanfordAILab · 2026-09-29
- RL Gains May Hit a Wall: Verifiable Data Limits and Stalling nanochart Benchmarks — QuintinPope5 · 2026-09-29
- TALES Benchmark on LLM Game-Playing Accepted to NeurIPS 2026, New Results Coming — tw_killian · 2026-09-29