SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Yue Zhang · hf · 2026-08-07

Understanding 3D scenes requires joint reasoning over heterogeneous information from multiple modalities, such as visual and geometric cues. However, existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, ignoring query-dependent needs. This introduces semantic noise and wastes computation.

To address this, the paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for 3D scene understanding. It incorporates two core modules:

Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks and attains highly competitive results on RGB-only video understanding benchmarks.

Original post →

More from Multimodal

Multimodal channel →