Switch language한국어
Back to the list

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

TL;DR AI

Key summary

2 min read
  1. OmniHuMo is a new large-scale multimodal motion dataset with over 5,000 hours and 3.2 million aligned sequences.

  2. The paper also presents AnyMo, a unified framework for generating human motion from arbitrary combinations of inputs.

  3. AnyMo combines a Residual FSQ motion tokenizer with a scalable masked modeling transformer.

  4. It can condition on text, speech, music, and trajectory signals in one system.

  5. This broadens motion generation control for animation, robotics, and multimodal interaction.

Read the original