OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

OMG: к генерации движений с открытым словарём с помощью смеси контроллеров
Xin Chen, Han Liang, Lan Xu, Sibei Yang, Ruonan Zhang, Sihan Ren, Yuecheng Xu, Jingyi Yu, Jiacheng Bao, Ruichi Zhang
2024-06-16

CLIP token embeddings alignmentMixture-of-Controllers (MoC) blockcross-attention with text-token-specific expertslarge-scale unlabeled motion dataset (20M instances)motion ControlNetopen-vocabulary text-to-motion generationunconditional diffusion model (1B parameters)zero-shot text-to-motion
We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods of-ten fail or produce implausible motions with unseen text inputs, which limits the applications. In this paper, we present OMG, a novel framework, which enables compelling motion generation from zero-shot open-vocabulary text prompts. Our key idea is to carefully tailor the pretrain-then-finetune paradigm into the text-to-motion generation. At the pre-training stage, our model improves the gener-ation ability by learning the rich out-of-domain inherent motion traits. To this end, we scale up a large unconditional diffusion model up to 1B parameters, so as to utilize the massive unlabeled motion data up to over 20M motion instances. At the subsequent fine-tuning stage, we intro-duce motion ControlNet, which incorporates text prompts as conditioning information, through a trainable copy of the pre-trained model and the proposed novel Mixture-of-Controllers (MoC) block. MoC block adaptively rec-ognizes various ranges of the sub-motions with a cross-attention mechanism and processes them separately with the text-token-specific experts. Such a design effectively aligns the CLIP token embeddings of text prompts to var-ious ranges of compact and expressive motion features. Ex-tensive experiments demonstrate that our OMG achieves significant improvements over the state-of-the-art meth-ods on zero-shot text-to-motion generation. Project page: https://tr3e.github.io/omg-page.
1
Extensive experiments show OMG significantly improves zero-shot text-to-motion generation over state-of-the-art methods.
2
Fine-tuning introduces Motion ControlNet: a trainable copy of the pretrained model combined with a novel Mixture-of-Controllers (MoC) block that conditions on text prompts.
3
OMG enables zero-shot open-vocabulary text-to-motion generation by tailoring a pretrain-then-finetune paradigm.
4
Pretraining uses a large unconditional diffusion model scaled up to 1B parameters trained on over 20M unlabeled motion instances to learn rich out-of-domain motion traits.
5
The MoC block uses cross-attention to recognize sub-motion ranges and processes them with text-token-specific experts, aligning CLIP token embeddings to compact expressive motion features.

Text-to-motion generation system (OMG framework) for zero-shot open-vocabulary text prompts

Improving zero-shot open-vocabulary motion generation by pretraining a large unconditional diffusion model on massive unlabeled motion data and fine-tuning via Motion ControlNet with a Mixture-of-Controllers block that aligns CLIP text token embeddings to compact expressive motion features

Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
24
Access Type
Author Information
Authors
Xin Chen
Han Liang
Lan Xu
Sibei Yang
Ruonan Zhang
Sihan Ren
Yuecheng Xu
Jingyi Yu
Jiacheng Bao
Ruichi Zhang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%