ROMA: a 7B model that lets robots understand objects by shaking, weighing and touching them
机器之心 · wechat · 2026-10-10
Researchers from Gaoling GeWu-Lab (RUC), BAAI and others released ROMA, a system for real-world object-centric multi-sensory active perception — letting robots fill gaps in physical understanding by actively interacting (shaking, lifting, pressing) rather than passively observing.
Three components: (1) ROMI-2K, a dataset of 2,000 everyday objects and contents across 6 basic physical interactions, with synchronized visual, audio, tactile and force data; (2) ROMABench, 2,100 scene-level active perception tasks across single-chain, multi-chain and intent-driven reasoning; (3) ROMA-7B, fine-tuned from Qwen2.5-Omni, which decides which object to grasp, what interaction to perform, which modality to attend to, and when to stop, closing the reason-interact-feedback loop.
ROMA-7B scores 72.9% on ROMABench, on par with GPT-6 Astra and ahead of GPT-5.4 and Gemini 3.5 Flash — and stronger on multi-chain tasks requiring continuous interaction. It holds 61.4% on 132 real-world open-ended questions. Dataset and code are open-sourced.
More from Embodied
- Dyna launches Taku semi-humanoid robot to remove the human babysitting bottleneck — JasonMa2020 · 2026-10-10
- First-ever humanoid deathmatch was absolutely wild — cixliv · 2026-10-10
- QWM on real ANYmal-D and Unitree Go1: cross-embodiment transfer with no fine-tuning or warm-up — breadli428 · 2026-10-10
- QWM method: morphology encoder, adaptive reward normalizer, and latent morphology conditioning — breadli428 · 2026-10-10
- QWM zero-shot transfer to unseen morphologies vs. model-free and morphology-expert baselines — breadli428 · 2026-10-10
- Atomic Machines exits stealth with AI-native Matter Compiler that builds micromachines from code — DeryaTR_ · 2026-10-10