What Limits Vision-and-Language Navigation?

TL;DR AI
2 min readKey summary
Researchers introduced StereoNav, a vision-language-action navigation system for real-world robot navigation.
The paper argues VLN breaks down under perceptual noise and vague instructions, revealing weak spatial grounding and poor robustness.
StereoNav combines persistent target-location priors with stereo depth cues to improve navigation under blur, lighting changes, and underspecified commands.
Tests on benchmark datasets and real robot deployments show better accuracy and consistency than scaling-heavy prior methods.
