GEM: Generative Supervision Helps Embodied Intelligence

TL;DR AI
2 min readKey summary
Researchers introduced GEM, a generative-supervised embodied vision-language model that adds depth-map generation to pretraining.
The team also released GEM-4M, a dataset with grounding, reasoning, planning, and depth supervision for embodied learning.
GEM achieved state-of-the-art results on multiple embodied benchmarks.
A GEM-based vision-language-action model, GEM-VLA, showed better simulation and real-world task execution, suggesting depth supervision improves spatial understanding and robot control.
