VLM3: Vision Language Models Are Native 3D Learners
TL;DR AI
2 min readKey summary
VLM3 shows standard vision-language models can be adapted for multiple 3D tasks.
With focal-length unification, text-based pixel references, and scaled data mixtures, they handle depth estimation, pixel correspondence, camera pose, and object-level 3D understanding.
The system matches specialized vision models while requiring fewer task-specific changes and simpler training pipelines.
