Welcome.AIWelcome.AI
    Skip to content
    research

    ROMA: Advanced Multi-Sensory Perception System for Robots

    Recent research presents a significant advancement in how robotic systems perceive and interact with their environment. The study introduces ROMA, a system designed for Real-World Object-Centric Multi...

    arxiv.org•October 7, 2026•3 min read

    Key Facts

    • Implement ROMA technology to enhance robotic decision-making in complex environments.
    • Leverage the ROMI-2K dataset for training more effective multi-sensory robotics applications.
    • Foster proactive information-gathering capabilities to improve real-time operational efficiency.
    • Integrate large language models to advance robotic perception and interaction strategies.
    • Optimize feedback loops in robotic systems for superior adaptability and performance.

    Summary

    Paper: ROMA: LLM System for Real-World Object-Centric Multi-Sensory Active Perception

    Authors: Ruoxuan Feng, Yutong Chen, Ruihua Song, Huan Yang, Zhongyuan Wang, Guocai Yao, Di Hu

    Executive Summary

    Recent research presents a significant advancement in how robotic systems perceive and interact with their environment. The study introduces ROMA, a system designed for Real-World Object-Centric Multi-Sensory Active Perception, which employs a large language model (LLM) to enhance robots' capabilities in gathering and processing sensory information.

    Traditional multi-sensory robotic systems primarily focus on integrating available sensory data rather than actively seeking out missing information. This limitation can hinder their effectiveness in complex environments where decisive action is required. ROMA addresses this by establishing a reasoning-interaction-feedback loop, which allows the robot to identify what information is lacking and how to gather it. This proactive approach could significantly improve decision-making in real-time scenarios.

    To support ROMA's functionality, the researchers created the ROMI-2K dataset, consisting of nearly 2,000 objects and six types of interactions, all with synchronized sensory feedback. This large-scale dataset is crucial for training the model, enabling it to assess whether enough evidence has been gathered before making decisions. The training framework is structured in two stages, focusing on aligning different sensory modalities and enhancing the LLM’s ability to select the most informative interactions based on the feedback received.

    The concept of active perception is further developed in this research, described as perception chains. These chains illustrate how evidence collected during one interaction informs subsequent actions and reasoning processes. To evaluate the effectiveness of ROMA, the researchers established ROMA Bench, a benchmarking tool designed to assess the system's ability to handle tasks that require long-horizon planning and multi-attribute perception.

    Experimental results indicate that ROMA can successfully acquire missing evidence and tackle complex multi-sensory perception tasks that have posed challenges for existing methods. This capability lays a strong foundation for the development of active multi-sensory embodied agents, potentially transforming various applications in fields such as autonomous vehicles, robotic assistants, and industrial automation.

    The implications of this research are noteworthy. By enabling robots to actively engage with their environment and seek out necessary information, organizations could enhance their operational efficiency and decision-making processes. While these findings stem from controlled experiments rather than real-world deployments, they suggest promising avenues for future applications in robotic systems where adaptability and responsiveness are critical.

    Academic Abstract

    Humans inherently understand the physical world through an active process. When sensory evidence is insufficient to infer physical properties, we naturally interact with the environment by deciding what information is missing, how to acquire it, and when sufficient evidence has been obtained. In stark contrast, existing multi-sensory robot systems mainly integrate sensory inputs rather than actively acquiring missing evidence through interactions. In this work, we introduce ROMA, an LLM-based system for Real-World Object-Centric Multi-Sensory Active Perception. ROMA integrates vision, audio, tactile, and force sensing into a reasoning-interaction-feedback loop. The model identifies missing evidence and determines the target objects, interactions, and modalities, while a physical interface executes the selected interactions and collects the multi-sensory feedback. To support this capability, we construct ROMI-2K, a large-scale real-world multi-sensory object interaction dataset covering nearly 2,000 objects and 6 atomic interactions with synchronized sensory feedback. Building on these data, we develop a two-stage training framework that aligns sensory modalities and equips the LLM to assess evidence sufficiency, select informative interactions, and reason over the multi-sensory feedback. We further characterize active perception as perception chains, where acquired evidence guides subsequent interactions and reasoning, and establish ROMA Bench to evaluate single-attribute, long-horizon multi-attribute, and intent-driven active perception. Experiments show that ROMA can actively acquire missing evidence and solve complex, long-chain multi-sensory perception tasks that existing methods struggle to handle, laying a strong perceptual foundation for active multi-sensory embodied agents.

    Frequently Asked Questions

    What business problems does the ROMA system solve?

    The ROMA system addresses the limitations of traditional multi-sensory robotic systems by enabling robots to actively seek out missing sensory information, which could enhance decision-making in complex environments and improve operational efficiency.

    Which industries may benefit most from the implementation of ROMA?

    Industries such as logistics, manufacturing, and healthcare could benefit significantly from ROMA, as these sectors often require advanced robotic systems for tasks involving real-time decision-making and interaction with diverse objects in dynamic settings.

    What are the practical implementation considerations for deploying ROMA in a business?

    Practical implementation considerations for deploying ROMA may include the integration of the system with existing robotic platforms, ensuring compatibility with various sensors, and training personnel to operate and maintain the system effectively.

    What resources or expertise are needed to implement the ROMA system successfully?

    Successful implementation of the ROMA system may require expertise in robotics, machine learning, and data processing, as well as access to the ROMI-2K dataset for training and enhancing the system's capabilities.

    What competitive advantages could businesses gain from using the ROMA system?

    Businesses that utilize the ROMA system could gain competitive advantages such as improved efficiency in operations, enhanced adaptability to complex environments, and the ability to make more informed decisions in real-time, potentially leading to increased productivity and better customer service.

    Where AI Leaders Stay Informed

    The latest AI intelligence, case studies, and research — delivered to your inbox every week.

    Free to read. Unsubscribe anytime.