GEM: Generative Supervision Helps Embodied Intelligence
TL;DR AI
2 min readKey summary
Researchers introduced GEM, a vision-language model that adds depth map generation during pretraining to better learn spatial and physical cues for robotics.
They also released GEM-4M, a dataset with grounding, reasoning, planning, and depth supervision data to support embodied learning.
GEM and GEM-VLA achieved state-of-the-art results on embodied benchmarks and showed strong performance in both real-world and simulated robot tasks.
The work highlights how adding spatial supervision to vision-language pretraining can close a key gap in robotics.
