NexaRob News/Report and research

How does stereo vision improve humanoid control? The new EATR-Stereo framework

The new EATR-Stereo framework, developed by a research team from Peking University, improves humanoid control in long-term tasks through intelligent transfer of data from two cameras mounted on the head. In tests on a physical robot, it achieved 60% success in a complete task and 80% effectiveness with strong asymmetrical occlusion - which means

On this page

How does EATR-Stereo solve the stereo vision problem in humanoids?

Conventional vision systems in robots often lose values from the second view or combine both images without considering their role in the context of main processing. EATR-Stereo introduces an innovative approach: it preserves basic visual tokens from one camera while creating auxiliary tokens based on the second view - so-called Cross-View Auxiliary Tokens (CVAT). This additional data is not simply added but transmitted consciously, taking into account the robot's body configuration. This allows for a better understanding of space and more precise control of movements.

The key to this approach is not only image analysis but also the integration of proprioceptive data - information about the robot's body state. The framework uses segment-based body encoding, which adjusts how additional data is used depending on the current position and movement history. This allows the robot to process information in a more natural and adaptive way, rather than treating all situations as identical.

Tests on a physical humanoid: effectiveness in real-world conditions

The research was conducted on a physical humanoid with 33 degrees of freedom and 37 proprioceptive states, performing "search, approach, grasp, place, return" tasks lasting over 100 seconds. Under these conditions, EATR-Stereo achieved 60.0% success in the full task and 100.0% success in grasping an object - indicating high precision in key stages of operation.

An important test was a situation with strong asymmetric occlusion, where one camera could not see the object. Under these conditions, EATR-Stereo improved its ability to recover to 80%, while the CVAT method alone achieved only 30%. This shows that intelligent transmission of data from the second view has a real impact on operational stability in difficult conditions.

Why are preserving basic tokens and proprioceptive routing crucial?

Ablation studies have shown that not all elements of the framework are equally important. In particular, it was found that maintaining basic visual tokens is essential - their loss leads to a loss of spatial context. Also, combining inter-view features with proprioceptive data-driven structural routing proved crucial for efficiency.

This means that EATR-Stereo is not just additional image processing. It is a system that understands when and how to use information from the second view - depending on where the robot is, how it moves, and what its physical limitations are. This approach is important for long-term control, where errors accumulate faster than in short tasks.

Significance and limitations: what does this mean for the future of robotics?

EATR-Stereo shows that the development of visual-linguistic systems for humanoids is not just about more data, but about its intelligent processing. The flow of information must take into account both the robot's physical limitations and the spatial context.

However, it is worth emphasizing that the results apply to one specific test model - a humanoid with 33 degrees of freedom. There is no information about scalability to other types of robots, nor about real-time performance. Additionally, the framework has not yet been implemented in commercial systems and has not been verified by independent research. Its value lies in demonstrating a new approach to integrating sensory data that can be used in future solutions. types of robots Additionally, the framework has not yet been implemented in commercial systems and has not been verified by independent research. Its value lies in demonstrating a new approach to integrating sensory data that can be used in future solutions.

Further context

Related NexaRob pages

Solutions, technologies and materials from NexaRob related to the topic of this article.

Share the material
English (United States)EN-US