ROMA: a 7B model that lets robots understand objects by shaking, weighing and touching them

机器之心 · wechat · 2026-10-10

Researchers from Gaoling GeWu-Lab (RUC), BAAI and others released ROMA, a system for real-world object-centric multi-sensory active perception — letting robots fill gaps in physical understanding by actively interacting (shaking, lifting, pressing) rather than passively observing.

Three components: (1) ROMI-2K, a dataset of 2,000 everyday objects and contents across 6 basic physical interactions, with synchronized visual, audio, tactile and force data; (2) ROMABench, 2,100 scene-level active perception tasks across single-chain, multi-chain and intent-driven reasoning; (3) ROMA-7B, fine-tuned from Qwen2.5-Omni, which decides which object to grasp, what interaction to perform, which modality to attend to, and when to stop, closing the reason-interact-feedback loop.

ROMA-7B scores 72.9% on ROMABench, on par with GPT-6 Astra and ahead of GPT-5.4 and Gemini 3.5 Flash — and stronger on multi-chain tasks requiring continuous interaction. It holds 61.4% on 132 real-world open-ended questions. Dataset and code are open-sourced.

Original post →

More from Embodied

Embodied channel →