SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Yue Zhang · hf · 2026-08-07
Understanding 3D scenes requires joint reasoning over heterogeneous information from multiple modalities, such as visual and geometric cues. However, existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, ignoring query-dependent needs. This introduces semantic noise and wastes computation.
To address this, the paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for 3D scene understanding. It incorporates two core modules:
- SMART (Semantic-guided Modality Adaptive Routing): Selects task-relevant modalities using semantic priors, text-modality alignment, and modality quality.
- MAGE (Modality-Aware Gating Expert): Leverages modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning.
Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks and attains highly competitive results on RGB-only video understanding benchmarks.
More from Multimodal
- MiniMax H3 Forgets Lyrics, Generates Gibberish Mimicking English Like the 1972 Hit — cocktailpeanut · 2026-08-07
- ByteDance's Dreamina Models Available to US Enterprises via fal Platform — gorkem · 2026-08-07
- AI Music Video Case Study: Generating a Full MV Using Grok — bennash · 2026-08-07
- ByteDance Seedream Layerize API Splits Images into Editable Transparent Layers — OdinLovis · 2026-08-07
- MiniMax H3 Preview: 2K Resolution and 5× Turbo Speed Incoming — RobbaW · 2026-08-07
- MiniMax H3 Outperforms Seedance 2.5 in Video Prompt Adherence — ZabihullahAtal · 2026-08-07