ROMA: Teaching Robots to Actively Explore and Understand Objects

Researchers from Renmin University and Beijing Academy of Artificial Intelligence introduce ROMA, a system that equips robots with multisensory active perception, enabling them to autonomously decide which object to interact with, how to interact, and which modalities to use, forming a reasoning-interaction-feedback loop that matches GPT-6 Astra performance on real-world tasks.

Machine Heart
Machine Heart
Machine Heart
ROMA: Teaching Robots to Actively Explore and Understand Objects

Human perception of objects is not passive; we actively lift, shake, squeeze, and press to infer properties like weight, contents, and softness. The ROMA (Real-World Object-Centric Multi-Sensory Active Perception) system, developed by researchers from Renmin University's GeWu-Lab, the Beijing Academy of Artificial Intelligence (BAAI), and AresoX, aims to give robots this same active, multisensory exploration capability.

ROMI-2K: Large-Scale Real-World Multimodal Interaction Dataset

A core challenge is teaching robots what information different physical interactions yield. To address this, the team built ROMI-2K , a dataset covering nearly 2,000 everyday objects and object-content combinations . Data was collected using both handheld devices and a real tabletop robotic arm across six fundamental interactions : lifting, pressing, colliding, pinching, shaking, and rotating. Each interaction synchronously captures four modalities : vision, audio, touch, and force. Physical attributes such as material, hardness, roughness, texture, contents, and weight are densely annotated. On top of this, the researchers constructed scene-level multimodal active perception QA data to train models to select appropriate interactions and modalities for a given task.

ROMA Bench: From Single Interaction to Perception Chains

In active perception, one interaction's observation influences the next. For example, shaking an object to detect contents: if a sound is heard, the object is likely not empty; if silent, it might be filled solid, requiring a lift to gauge weight. This creates a perception chain of reasoning-interaction-feedback. Based on this concept, the team built ROMA Bench on the real-tabletop subset of ROMI-2K, containing 2,100 scene-level active perception tasks across three categories: (1) Single-Chain (single-attribute reasoning), (2) Multi-Chain (multi-attribute, long-chain reasoning), and (3) Intent-Driven (inferring required attributes from implicit user intent). Tasks involve properties like hardness, roughness, texture, contents, material, and weight. This benchmark moves beyond the simple "interact once then answer" paradigm to a full loop of goal selection, action selection, modality selection, and continuous reasoning.

ROMA System: Physical Interface and the ROMA-7B Model

The ROMA system connects vision, audio, touch, and force sensors to a robotic arm via a unified physical interaction interface, allowing the model to output real-world interaction commands. The six base interactions from ROMI-2K plus grasping are proceduralized. Using the AnyGrasp grasping policy, enhanced with point-cloud completion, grasp filtering, and augmentation mechanisms , the system achieves stable real-world grasping, reliably translating interaction decisions into physical actions.

On this hardware, the team trained ROMA-7B based on Qwen 2.5-Omni through multi-sensor joint alignment and active-perception scenario fine-tuning . ROMA-7B can, given a task goal and current observations, actively decide which object to grasp, which interaction to perform, which modalities to attend to, and when to stop exploring . It dynamically adjusts subsequent interactions based on newly acquired multimodal feedback, closing the "reasoning-interaction-feedback-reasoning" loop.

Benchmark Results: Matching GPT-6 Astra, Leading Other Frontier Models

On the 2,100 ROMA Bench tasks (multiple-choice, true/false, ranking), ROMA-7B achieves a 72.9% overall success rate , on par with the latest GPT-6 Astra and clearly outperforming GPT-5.4 and Gemini 3.5 Flash . GPT-6 Astra excels at single-step, visually inferable attributes due to stronger visual understanding and general reasoning. ROMA-7B shows stronger advantage on Multi-Chain tasks requiring continuous interaction and fusion of multiple modalities, especially audio , surpassing GPT-6 Astra there. Untuned same-size baselines Qwen 2.5-Omni and Qwen 3-Omni perform poorly on ROMA Bench's goal selection, instruction following, and multimodal active reasoning, highlighting the critical value of ROMI-2K for learning active perception.

In real-world experiments across 8 multi-object scenes and 132 open-ended QA questions , ROMA-7B maintains a 61.4% overall success rate , again close to GPT-6 Astra and far ahead of other frontier models, demonstrating that the active perception skills learned from ROMI-2K transfer to physical deployment, not just benchmark formats.

Future Directions

ROMA currently focuses on six basic interactions as a starting point. As a foundational system, it can inspire further work on: richer interaction skills, unification of active perception and manipulation policies, multi-turn dialogue and long-horizon reasoning, persistent cross-scene object memory, and active perception for diverse robot morphologies. The ultimate goal remains: "I see, I touch, I understand" — enabling robots to actively interact with the world and, through repeated exploration, truly understand it.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Embodied AIRoboticsROMAmultimodal perceptionactive perceptionQwen 2.5-OmniROMA BenchROMI-2K
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.