FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

TL;DR AI
2 min readKey summary
Researchers introduced FineVLA, an open framework for training vision-language-action robot policies with fine-grained execution instructions.
It combines multiple robot datasets, a held-out benchmark, and a specialized annotator to create detailed supervision for actions and contact.
Policies trained with mixed fine-grained and goal-level instructions outperformed raw goal-only training.
The approach improved controllability over pose, color, and approach direction while preserving or improving task success.
