Qwen Releases OmniReasoning: Audio-Visual Joint Reasoning Benchmark and 30B Model
Qwen · hf · 2026-10-06
Qwen introduces a full stack for audio-visual joint reasoning, arguing existing omni-modal benchmarks and methods treat modalities independently.
Deliverables
- OmniReasoningBench: 1,150 MCQ and open-ended questions where audio and visual evidence are both indispensable, spanning in-video and beyond-video reasoning.
- OmniQA data engine: auto-builds evidence-grounded QA pairs with time-stamped clue chains, yielding 112K SFT and 19K RL samples.
- MFSD: Modality-Factored Self-Distillation for token-level credit assignment across modal clues.
Results: OmniReasoning-30B-A3B scores 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, +12.8 and +9.3pp over Qwen3-Omni-30B-A3B-Thinking, with gains on Video-MME-v2 and other benchmarks.
More from Multimodal
- Open-source huashu-art-motion skill animates 35 art styles via coding agents — AlchainHust · 2026-10-06
- AI-generated anime stage clip wows Reddit with smoke and lightning vibes — Amazing_Skill_6080 · 2026-10-06
- Creator shares 120-prompt workflow building an AI video game with Seedance 2.5, Godot and Magnific — techhalla · 2026-10-06
- TikTok 5.6B-Video Dataset Trending on Hugging Face — datasocial · 2026-10-06
- DEPICT: Training-Free Alignment Metric Boosts Negation Accuracy from 19% to 88% — swordhealth · 2026-10-06
- IDU: Unified Unlearning and One-Step Distillation for Flow and Diffusion Models — Aleksei Leonov · 2026-10-06