HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation

TL;DR AI
2 min readKey summary
Researchers introduced HCSG, a human-centric vision-language navigation framework for robots in indoor environments.
It combines a human understanding module, VLM-based intention inference, and topological map planning to predict pedestrian motion and guide navigation.
A social distance loss further improves socially compliant behavior and reduces collisions.
On HA-VLNCE, HCSG reports a 14% higher success rate and 34% lower collision rate than prior methods.
