OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers
OMG: к генерации движений с открытым словарём с помощью смеси контроллеров
2024-06-16
SCID: 54.1/bexsnujn
Discuss with AI
CLIP token embeddings alignmentMixture-of-Controllers (MoC) blockcross-attention with text-token-specific expertslarge-scale unlabeled motion dataset (20M instances)motion ControlNetopen-vocabulary text-to-motion generationunconditional diffusion model (1B parameters)zero-shot text-to-motion
Figures from the paper
Abstract (AI)
We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods of-ten fail or produce implausible motions with unseen text inputs, which limits the applications. In this paper, we present OMG, a novel framework, which enables compelling motion generation from zero-shot open-vocabulary text prompts. Our key idea is to carefully tailor the pretrain-then-finetune paradigm into the text-to-motion generation. At the pre-training stage, our model improves the gener-ation ability by learning the rich out-of-domain inherent motion traits. To this end, we scale up a large unconditional diffusion model up to 1B parameters, so as to utilize the massive unlabeled motion data up to over 20M motion instances. At the subsequent fine-tuning stage, we intro-duce motion ControlNet, which incorporates text prompts as conditioning information, through a trainable copy of the pre-trained model and the proposed novel Mixture-of-Controllers (MoC) block. MoC block adaptively rec-ognizes various ranges of the sub-motions with a cross-attention mechanism and processes them separately with the text-token-specific experts. Such a design effectively aligns the CLIP token embeddings of text prompts to var-ious ranges of compact and expressive motion features. Ex-tensive experiments demonstrate that our OMG achieves significant improvements over the state-of-the-art meth-ods on zero-shot text-to-motion generation. Project page: https://tr3e.github.io/omg-page.
Key Findings
1
Extensive experiments show OMG significantly improves zero-shot text-to-motion generation over state-of-the-art methods.
2
Fine-tuning introduces Motion ControlNet: a trainable copy of the pretrained model combined with a novel Mixture-of-Controllers (MoC) block that conditions on text prompts.
3
OMG enables zero-shot open-vocabulary text-to-motion generation by tailoring a pretrain-then-finetune paradigm.
4
Pretraining uses a large unconditional diffusion model scaled up to 1B parameters trained on over 20M unlabeled motion instances to learn rich out-of-domain motion traits.
5
The MoC block uses cross-attention to recognize sub-motion ranges and processes them with text-token-specific experts, aligning CLIP token embeddings to compact expressive motion features.
Research Object
Text-to-motion generation system (OMG framework) for zero-shot open-vocabulary text prompts
Research Subject
Improving zero-shot open-vocabulary motion generation by pretraining a large unconditional diffusion model on massive unlabeled motion data and fine-tuning via Motion ControlNet with a Mixture-of-Controllers block that aligns CLIP text token embeddings to compact expressive motion features
Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
24
Access Type
Author Information
Download PDF
Subscribe to digest