Diagnosing Compositional Generalization in Sequential Robot Tasks

TL;DR AI
2 min readKey summary
Researchers examined why sequential robot policies fail on unseen instruction combinations.
They decomposed the generalization gap into marginal instruction shift, instruction-compositional shift, and context-action shift.
Full enumeration of all instruction tuples is unnecessary; a structured subset that covers action-relevant dependencies can preserve strong out-of-distribution performance.
Sparse training often fails due to instruction steering effects rather than missing low-level skills, and finetuning with one demo per task sharply improves success.
