AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
TL;DR AI
2 min readKey summary
OmniHuMo is a new large-scale multimodal motion dataset with over 5,000 hours and 3.2 million aligned sequences.
The paper also presents AnyMo, a unified framework for generating human motion from arbitrary combinations of inputs.
AnyMo combines a Residual FSQ motion tokenizer with a scalable masked modeling transformer.
It can condition on text, speech, music, and trajectory signals in one system.
This broadens motion generation control for animation, robotics, and multimodal interaction.
