OmniReasoning: audio-visual joint reasoning benchmark lifts Qwen3-Omni by 12.8 points
_akhaliq · x · 2026-10-07
OmniReasoning targets the gap in audio-visual joint reasoning, where existing benchmarks and training treat modalities independently. The release includes:
- OmniReasoningBench: 1,150 multiple-choice and open-ended questions where both audio and visual evidence are indispensable, spanning reasoning over and beyond video.
- OmniQA data engine: automatically builds evidence-grounded QA pairs with timestamped clue chains, yielding OmniReasoning-SFT-112K and RL-19K training sets.
- Modality-Factored Self-Distillation (MFSD): evaluates sampled responses under modality-specific clue contexts for token-level credit assignment.
OmniReasoning-30B-A3B scores 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving base Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 points, with gains on general and long-video benchmarks like Video-MME-v2.
More from Multimodal
- Mirage's Tesseract + Opus 5.5 generated its entire launch video, exported to After Effects — aziz4ai · 2026-10-07
- Nano Banana 2.1 vs Midjourney 8.2: same-prompt image comparison surfaces — miilesus · 2026-10-07
- A lot of AI slop films are just porn, observer points out — moonsandhues · 2026-10-07
- First AI-Made Film Gets Traditional Theatrical Run, Submits for Best Animated Feature Oscar — Uncanny_Harry · 2026-10-07
- AI turns kids' doodles into fluffy living creatures — horned tiger included — anthara_ai · 2026-10-07
- Seedance 2.5 AI video of Dragon Ball battle wows with detail and motion — anthara_ai · 2026-10-07