Switch language한국어
Back to the list

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced WorldVLN, an autoregressive world-action model for aerial vision-language navigation.

  2. The system predicts short-horizon state changes from video, decodes them into waypoint actions, and updates context in a closed loop.

  3. It uses a two-stage training pipeline: navigation-grounded pretraining plus Action-aware GRPO reinforcement learning.

  4. WorldVLN reports more than 12% success-rate gains over baselines on indoor and outdoor benchmarks and transfers zero-shot to real drone flight.

Read the original