Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions

TL;DR AI
2 min readKey summary
Researchers introduced an orientation-aware zero-shot action recognition method that learns multiview motion features and matches them with action text prompts at inference.
The approach improves recognition under viewpoint and body-orientation changes, a major source of domain shift in real-world video understanding.
It delivers stronger zero-shot results across benchmarks such as NTU-RGB+D, BABEL, and NW-UCLA, including surveillance-style settings.
The method also transfers better on seen actions, showing broader robustness without heavy manual labeling.
