NexaRob News/Report and research

How does stereoscopic vision improve humanoid control? The new EATR-Stereo framework

The new EATR-Stereo framework, developed by a research team from Peking University, improves humanoid control in long-term tasks through intelligent transfer of data from two cameras mounted on the head. In tests on a physical robot, it achieved 60% success in the full task and 80% effectiveness with strong asymmetrical occlusion - which means

On this page

How does EATR-Stereo solve the problem of stereoscopic vision in humanoids?

Conventional computer vision systems in robots often lose values from the second view or combine both images without considering their role in the context of main processing. EATR-Stereo introduces an innovative approach: it preserves basic visual tokens from one camera while creating auxiliary tokens based on the second view - so-called Cross-View Auxiliary Tokens (CVAT). This additional data is not simply added, but transferred in a conscious way, taking into account the robot's body configuration. This allows for a better understanding of space and more accurate control of movements.

The key to this approach is not only image analysis, but also the integration of proprioceptive data - information about the state of the robot's body. The framework uses segment-based body encoding, which adjusts the way additional data is used depending on the current position and movement history. This allows the robot to not treat all situations as identical, but to process information in a more natural and adaptive way.

Tests on a physical humanoid: efficiency in real-world conditions

The research was conducted on a physical humanoid with 33 degrees of freedom and 37 proprioceptive states, performing "search, approach, grasp, place, return" tasks lasting over 100 seconds. Under these conditions, EATR-Stereo achieved 60.0% success in the full task and 100.0% success in grasping an object - which indicates high precision in key stages of operation.

An important test was a situation with strong asymmetrical occlusion, where one camera could not see the object. Under these conditions, EATR-Stereo improved the ability to recover to 80%, while the CVAT method alone achieved only 30%. This shows that intelligent transfer of data from the second view has a real impact on operational stability in difficult conditions.

Why are preserving basic tokens and proprioceptive routing crucial?

Ablation studies have shown that not all elements of the framework are equally important. In particular, it has been found that preserving basic visual tokens is essential - their loss leads to a loss of spatial context. Also, combining inter-view features with structural routing based on proprioceptive data proved crucial for efficiency.

This means that EATR-Stereo is not just an additional image processing step. It is a system that understands when and how to use information from the second view - depending on where the robot is located, how it moves, and what its physical limitations are. This approach is important for long-term control, where errors accumulate faster than in short tasks.

Significance and limitations: what does this mean for the future of robotics?

EATR-Stereo shows that the development of visual-linguistic systems for humanoids is not just about more data, but about its intelligent processing. The flow of information must take into account both the physical limitations of the robot and the spatial context.

However, it is worth emphasizing that the results apply to one specific test model - a humanoid with 33 degrees of freedom. There is no information on scalability to other types of robots or on real-time performance. In addition, the framework has not yet been implemented in commercial systems and has not been verified by independent research. Its value lies in demonstrating a new approach to integrating sensory data that can be used in future solutions. types of robots or its performance in real time. Furthermore, the framework has not yet been deployed in commercial systems and has not been verified by independent research. Its value lies in demonstrating a new approach to integrating sensory data that can be used in future solutions.

More context

Related NexaRob pages

Solutions, technologies and materials from NexaRob related to the subject of this article.

Share the material
EnglishEN