GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

TL;DR AI
2 min readKey summary
Researchers introduced GenPrior, a zero-shot skeleton action recognition framework that transfers knowledge from pre-trained text-to-motion models.
It uses dispersion-gated feature fusion to inject structural motion cues into text embeddings, plus generative prototype refinement to better fit unseen actions.
GenPrior tackles the semantic-kinematic gap that weakens text-prototype methods and improves recognition of unseen skeleton actions.
The method outperforms prior approaches on NTU-60, NTU-120, and PKU-MMD in both zero-shot and generalized zero-shot settings.
