NexaRob News/Report and research

Vision-based navigation for quadruped robots: mapless, precise, and with deep control.

Researchers at Amazon Science have developed a new approach to navigation for quadruped robots that allows them to safely navigate complex terrains and perform inspections based solely on images from a single camera. In tests, they achieved high efficiency, even with large viewing angles, without the need for depth measurement or precise maps.

On this page

Instead of maps - images and intelligence.

Many autonomous navigation systems for robots rely on precise maps or depth sensors, which limits their application in unpredictable environments. The new approach developed by Amazon Science teams eliminates this dependency - it works solely based on data from a single RGB camera. This means that the robot can move and perform inspections in areas without pre-prepared maps, opening up new possibilities in logistics, infrastructure monitoring, or rescue operations.

The key to this solution is not only the image but also how it is interpreted. Instead of providing the robot with a ready-made target image, the system continuously updates the target representation based on the video stream from the camera. This allows for better resilience to poor detection quality or distance to the object - even if the target is not initially clearly visible, the robot can "refine" it during movement.

As part of this approach, the team introduced real-time target reflection, based on object detection in the video stream. Each new frame updates the target representation, allowing for smooth adaptation to changing conditions and increasing resilience to detection interference.

Phases and control: how does the robot understand when to stop?

The system divides the navigation process into two phases - long-distance approach and close orbital inspection. Each phase has its own goals and requires a different control strategy. Long distance requires a stable direction, while close inspection requires precise maintenance of the object in the center of the field of view. To achieve this, the team designed a policy structure with separate controls that use the same basic navigation architecture but are adapted to the specifics of each phase.

This allows for smooth transitions between phases without the need for complete reconfiguration. During testing, the system proved particularly effective at large viewing angles, which means that the robot can approach the target from the side or at a large angle - a situation in which traditional methods often fail.

Separating the phases allows for optimizing control: in the long-distance phase, the robot focuses on direction and stability of movement, while in the close-up phase - on precise positioning and keeping the object in the center of the frame.

Stopping without depth perception: how does the vision mechanism work?

visual stopping mechanism
Stopping without depth perception: how does the vision mechanism work - illustrative visualization

One of the biggest challenges in mapless navigation is not only finding the target, but also stopping accurately at it. Many systems rely on depth sensors, which are expensive and prone to interference. The new halting mechanism works purely visually: it analyzes the proportion of the detection rectangle (bbox) to the total image area and applies smooth temporal processing to eliminate vibrations caused by robot movement.

In addition, the system uses a visual servoing loop, which keeps the object in the center of the frame even during unstable movement or camera vibrations caused by walking. This allows for stable stopping at a distance of 1.0-1.2 meters from the target - which is crucial for precise inspection.

This mechanism does not require depth measurement or additional sensors. Instead, it uses dynamic changes in image proportions as a signal to stop, ensuring stability and repeatability of results.

Tests and generality: effectiveness without architectural limitations

In each case, the system achieved high final effectiveness. To verify independence from a specific model, tests were also carried out on two different architectures: ViNT and NoMaD - without any modifications. The results confirmed that the approach is universal and can be applied in various navigation systems.

This means that the technology is not tied to a specific robot model or algorithm - it can be implemented in existing or future solutions. However, the limitation remains the lack of data from real-world environments outside of test scenarios, which means that effectiveness in extreme conditions (e.g., rain, fog) has not yet been confirmed.

Research shows that the system works stably both with a simple trajectory and at a large viewing angle - which is crucial for applications in hard-to-reach areas.

Significance and future: from tests to practice

The new approach shows that autonomous navigation in difficult conditions does not have to rely on complex sensor systems. A vision-based approach with dynamic goal reflection and deep phase separation may be the key to developing robots in industry, logistics or rescue - where it is impossible to prepare maps or install additional sensors.

Although the system was presented as a research demonstration, its application can be extended to real-world scenarios. However, the lack of data on deployment in real conditions or long-term tests leaves open questions about resilience and stability over time.

In the future, it will be possible to develop this approach for more complex scenarios, such as cooperation between multiple robots or integration with logistics management systems.

Editorial transparency

Sources and reference materials

The article was prepared by NexaRob based on an analysis of available source materials. The following materials were used to verify information and expand the context.

1source material
0primary
1publicly shown
  1. Online sourceInformation verification

    Multi-phase vision-based navigation and inspection for legged robots with online goal refinement and vision-only halting

    Amazon Science - Roboticsamazon.science

How to read this section? Sources are materials used during research and verification. The article is an original NexaRob report, not a reprint of the indicated publications.

Share the material
Go to sources
EnglishEN