Robostral Navigate
TL;DR AI
2 min readKey summary
Researchers introduced Robostral Navigate, an 8B vision-language navigation model that uses only monocular RGB images to predict waypoints in image space.
Trained on 2.4 million simulated trajectories with prefix-caching and reinforcement learning, it improves training efficiency and cuts cost and time.
The model sets state-of-the-art results on R2R-CE and RxR-CE, showing strong indoor navigation performance without expensive sensors or maps.
Its RGB-only design makes navigation policies easier to deploy across different robot platforms and reduces recalibration overhead.
