Switch language한국어
Back to the list

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

TL;DR AI

Key summary

2 min read
  1. ShadowDancer is a new framework that turns demonstration videos into reusable action representations for frame-level control of video world models.

  2. It learns transferable dynamics by pairing videos with the same motion but different appearances, then predicting one view from the other through cross-shadow prediction.

  3. This lets ordinary videos serve as action supervision without action labels, motion estimation, or fine-tuning.

  4. The method outperforms strong baselines on action transfer and longer rollouts.

  5. Overall, it could broaden how interactive video models are trained and controlled across diverse environments.

Read the original